IP Library › Granted Patent US 11,556,706
Granted Patent B2
US 11,556,706 · App. 16/423,333 · Granted Jan 17, 2023

Effective retrieval of text data based on semantic attributes between morphemes

Inventors: Seiji Okura (Meguro, JP); Masahiro Kataoka (Kamakura, JP)
Assignee: FUJITSU LIMITED
G06F40/268G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,556,706
App. No.
16/423,333
Granted
Jan 17, 2023
Kind
B2
Abstract

An apparatus generates an index including positions of morphemes included in a target text data and semantic attributes between the morphemes corresponding to the positions. The apparatus gives information including positions of morphemes included in an input query and semantic attributes between the morphemes corresponding to the positions to the query, and executes a retrieval on the target text data, based on the information given to the query and the index.

Claims (32)

1. A method performed by a processor included in a retrieval apparatus, the method comprising:

generating an index using a syntactic analysis and semantic analysis, the index including positions of morphemes included in a target text data and semantic attributes between the morphemes corresponding to the positions;

giving information including first positions of first morphemes included in an input query and semantic attributes between the first morphemes corresponding to the first positions to the query; and

executing a retrieval on the target text data based on the information given to the query and the index from a storage device, wherein

the generating the index identifies, as an expression analogous to a composite word, a group of morphemes corresponding to three or fewer nodes of nodes directly connecting to a node of a morpheme of a composite word in a syntactic tree and a semantic structure acquired as a result of the syntactic analysis and the semantic analysis thereon.

2. The method of claim 1 , further comprising obtaining the index, wherein the executing a retrieval executes a retrieval on the target text data based on the obtained index and the information given to the query.

3. The method of claim 1 , wherein the semantic attributes between the morphemes are information indicating a morpheme being a starting point of a dependency between the morphemes and a morpheme being an end point of the dependency.

4. The method of claim 1 , wherein the target text data is a character string including two or more words having semantic attributes.

5. The method of claim 1 , wherein the executing a retrieval is based on whether or not a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the information given to the query agree with a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the index.

6. A non-transitory, computer-readable recording medium having stored therein a program for causing a computer included in a retrieval apparatus to execute a process comprising:

generating an index using a syntactic analysis and semantic analysis, the index including first positions of first morphemes included in a target text data and semantic attributes between the morphemes corresponding to the first positions;

giving information including positions of morphemes included in an input query and semantic attributes between the first morphemes corresponding to the positions to the query; and

executing a retrieval on the target text data based on the information given to the query and the index, wherein

the generating the index identifies, as an expression analogous to a composite word, a group of morphemes corresponding to three or fewer nodes of nodes directly connecting to a node of a morpheme of a composite word in a syntactic tree and a semantic structure acquired as a result of the syntactic analysis and the semantic analysis thereon.

7. The non-transitory, computer-readable recording medium of claim 6 , the process further comprising obtaining the index, wherein the executing a retrieval executes a retrieval on the target text data based on the obtained index and the information given to the query.

8. The non-transitory, computer-readable recording medium of claim 6 , wherein the semantic attributes between the morphemes are information indicating a morpheme being a starting point of a dependency between the morphemes and a morpheme being an end point of the dependency.

9. The non-transitory, computer-readable recording medium of claim 6 , wherein the target text data piece is a character string including two or more words having semantic attributes.

10. The non-transitory, computer-readable recording medium of claim 6 , wherein the executing a retrieval is based on whether or not a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the information given to the query agree with a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the index.

11. A retrieval apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

generate an index using a syntactic analysis and semantic analysis, the index including positions of morphemes included in a target text data and semantic attributes between the morphemes corresponding to the positions,

give information including first positions of first morphemes included in an input query and semantic attributes between the first morphemes corresponding to the first positions to the query, and

execute a retrieval on the target text data based on the information given to the query and the index, wherein

the generated index identifies, as an expression analogous to a composite word, a group of morphemes corresponding to three or fewer nodes of nodes directly connecting to a node of a morpheme of a composite word in a syntactic tree and a semantic structure acquired as a result of the syntactic analysis and the semantic analysis thereon.

12. The retrieval apparatus of claim 11 , wherein:

the processor is further configured to obtain the index; and

the processor execute a retrieval on the target text data based on the obtained index and the information given to the query.

13. The retrieval apparatus of claim 11 , wherein the semantic attributes between the morphemes are information indicating a morpheme being a starting point of a dependency between the morphemes and a morpheme being an end point of the dependency.

14. The retrieval apparatus of claim 11 , wherein the target text data is a character string including two or more words having semantic attributes.

15. The retrieval apparatus of claim 11 , wherein

the processor executes a retrieval, based on whether or not a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the information given to the query agree with a morpheme being a starting point of a dependency between morphemes and a morpheme being an end point of the dependency in the index.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2019
From: OKURA, SEIJI; KATAOKA, MASAHIRO
To: FUJITSU LIMITED
Reel/Frame 049290/0492 →
Priority Claims (1)
JP JP2018-106940 · Jun 4, 2018 · national
Continuity (1)
Related Publication 20190370328A1 · Dec 5, 2019