IP Library Granted Patent US 10,460,037
Granted Patent B2
US 10,460,037 · App. 15/606,162 · Granted Oct 29, 2019

Method and system of automatic generation of thesaurus

Inventor: Yury Grigorievich Zelenkov (Orekhovo-Zuevo, RU)
Assignee: YANDEX EUROPE AG
G06F17/2795G06F17/274G06F17/2705G06F17/277G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,037
App. No.
15/606,162
Granted
Oct 29, 2019
Kind
B2
Abstract

A method of automatic generation of a digital thesaurus, the method comprising: parsing the digital text and determining a first lexical unit and a second lexical unit; for each entry of the first lexical unit: selecting n-number of sequential units adjacent to the first lexical unit; generating a first context parameter for the first lexical unit, the first context parameter comprising an indication of each unit of the n-number of sequential units and a frequency of co-occurrence of each unit with the first lexical unit in the digital text; for each entry of the second lexical: selecting, n-number of sequential units adjacent to the second lexical unit; generating a second context parameter; determining a lexical unit relation parameter for the first lexical unit and the second lexical unit by: an interrelation analysis and an analysis of entry co-occurrence.

Claims (83)

1. A method of automatic generation of a digital thesaurus, the method executable by a server coupled to a semantic relationship database, the method comprising:

acquiring by the server, an indication of a digital text, the digital text comprising one or more sentences;

parsing, by the server, the digital text and determining a first lexical unit and a second lexical unit;

for each entry of the first lexical unit in the digital text:

selecting, by the server, n-number of sequential units adjacent to the first lexical unit;

generating a first plurality of unit-pairs, the first plurality of unit-pairs comprising the first lexical unit paired with each unit of the n-number of sequential units adjacent to the first lexical unit;

generating, by the server, a first context parameter for the first lexical unit, the first context parameter comprising a frequency of co-occurrence of each unit-pair of the first plurality of unit-pairs within the one or more sentences of the digital text;

for each entry of the second lexical unit in the digital text:

selecting, by the server, n-number of sequential units adjacent to the second lexical unit;

generating a second plurality of unit-pairs, the second plurality of unit-pairs comprising the second lexical unit paired with each unit of the n-number of sequential units adjacent to the second lexical unit;

generating, by the server, a second context parameter for the second lexical unit, the second context parameter comprising a frequency of co-occurrence of each unit-pair of the second plurality of unit-pairs within the one or more sentences of the digital text;

determining, by the server, a lexical unit relation parameter for the first lexical unit and the second lexical unit, the lexical unit relation parameter indicative of a semantic link between the first lexical unit and the second lexical unit, the lexical unit relation parameter being determined by:

an interrelation analysis of the first context parameter and the second context parameter, the interrelation analysis comprising:

determining a first inclusion parameter indicative of the inclusion of the first context parameter into the second context parameter;

determining a second inclusion parameter indicative of the inclusion of the second context parameter into the first context parameter;

determining a first similarity parameter between the first context parameter and the second context parameter;

an analysis of entry co-occurrence of the first lexical unit and the second lexical unit in the digital text, the analysis of entry co-occurrence comprising:

determining a co-occurrence parameter indicative of a frequency of the first lexical unit and the second lexical unit being contained within a same sentence of the digital text;

wherein, upon determination that the first inclusion parameter and the second inclusion parameter are below a first threshold, the lexical unit relation parameter is indicative of:

a synonymous relationship if the first similarity parameter is above a second threshold and the co-occurrence parameter is below a third threshold;

an antonymous relationship if the first similarity parameter is above a fourth threshold and the co-occurrence parameter is above a fifth threshold;

an associative link if the first similarity parameter is below a sixth threshold; and

storing, by the server, the lexical unit relation parameter in the semantic relationship database.

2. The method of claim 1 , further comprising associating a grammatical type to each word of the digital text before determining the first lexical unit and the second lexical unit.

3. The method of claim 2 , wherein the lexical unit is one of:

a word, the word being determined based on its associated grammatical type; and

a phrase, the phrase being a group of two or more words determined based on the associated grammatical type of one of the two or more words.

4. The method of claim 3 , further comprising lemmatizing the first and second lexical units and the words of the digital text before determining the frequency of co-occurrence.

5. The method of claim 1 , wherein the n-number of sequential units are at least one of sequentially preceding, sequentially following, or sequentially preceding and following the first and second lexical unit, respectively.

6. The method of claim 1 , wherein the analysis of entry co-occurrence comprises determining a co-occurrence parameter indicative of a frequency of the first lexical unit and the second lexical unit being contained within a same sentence of the digital text.

7. The method of claim 1 , wherein the lexical unit relation parameter for the first and second lexical unit is indicative of a hypernym-hyponym relationship if one of the first inclusion parameter or the second inclusion parameter is above a threshold.

8. The method of claim 6 , wherein the interrelation analysis further comprises:

further parsing the digital text, by the server, to determine a third lexical unit;

for each entry of the third lexical unit in the text:

selecting, by the server, n-number of sequential units adjacent to the third lexical unit;

generating a third plurality of unit-pairs, the third plurality of unit-pairs comprising the third lexical unit paired with each unit of the n-number of sequential units adjacent to the third lexical unit;

generating, by the server, the third context parameter for the third lexical unit, the third context parameter comprising a frequency of co-occurrence of each unit-pairs of the third plurality of unit-pairs within the one or more sentences of the digital text; and

determining a second similarity parameter of the third context parameter with the second context parameter.

9. The method of claim 8 , wherein the lexical unit relation parameter for the first, the second and third lexical unit is indicative of the holonym-meronym relationship if the first inclusion parameter and the second inclusion parameter is above a seventh threshold, and the second similarity parameter is below a eight threshold.

10. A server for automatic generation of a digital thesaurus, the server comprising:

a network interface for communicatively coupling to a communication network;

a processor coupled to the network interface, the professor configured to:

acquire by the server, an indication of a digital text, the digital text comprising one or more sentences;

parse, by the server, the digital text and determining a first lexical unit and a second lexical unit;

for each entry of the first lexical unit in the digital text:

select, by the server, n-number of sequential units adjacent to the first lexical unit;

generate a first plurality of unit-pairs, the first plurality of unit-pairs comprising the first lexical unit paired with each unit of the n-number of sequential units;

generate, by the server, a first context parameter for the first lexical unit, the first context parameter comprising a frequency of co-occurrence of each unit-pair within the digital text;

for each entry of the second lexical unit in the digital text:

select, by the server, n-number of sequential units adjacent to the second lexical unit;

generate a second plurality of unit-pairs, the second plurality of unit-pairs comprising the second lexical unit paired with each unit of the n-number of sequential units adjacent to the second lexical unit;

generate, by the server, a second context parameter for the second lexical unit, the second context parameter comprising a frequency of co-occurrence of each unit-pair within the one or more sentences of the digital text;

determine, by the server, a lexical unit relation parameter for the first lexical unit and the second lexical unit, the lexical unit relation parameter indicative of a semantic link between the first lexical unit and the second lexical unit, the lexical unit relation parameter being determined by:

an interrelation analysis of the first context parameter and the second context parameter, the interrelation analysis comprising:

determining a first inclusion parameter indicative of the inclusion of the first context parameter into the second context parameter;

determining a second inclusion parameter indicative of the inclusion of the second context parameter into the first context parameter;

determining a first similarity parameter between the first context parameter and the second context parameter;

an analysis of entry co-occurrence of the first lexical unit and the second lexical unit in the digital text, the analysis of entry co-occurrence comprising:

determining a co-occurrence parameter indicative of a frequency of the first lexical unit and the second lexical unit being contained within a same sentence of the digital text;

wherein, upon determination that the first inclusion parameter and the second inclusion parameter are below a first threshold, the lexical unit relation parameter is indicative of:

a synonymous relationship if the first similarity parameter is above a second threshold and the co-occurrence parameter is below a third threshold:

an antonymous relationship if the first similarity parameter is above a fourth threshold and the co-occurrence parameter is above a fifth threshold:

an associative link if the first similarity parameter is below a sixth threshold; and

store, by the server, the lexical unit relation parameter in the semantic relationship database.

11. The server of claim 10 , further configured to lemmatize the first and second lexical units and the words of the digital text before determining the frequency of co-occurrence.

12. The server of claim 10 , wherein the n-number of sequential units are at least one of sequentially preceding, sequentially following, or sequentially preceding and following the first and second lexical unit, respectively.

13. The server of claim 12 , wherein upon determining that the n-number of sequential units adjacent to a given occurrence of the first lexical unit spans into an additional sentence adjacent thereto, the processor is configured to generate a respective first context parameter associated with the given occurrence comprises using a subset of the n-number of sequential units, the subset being units from the sentence of the given occurrence.

14. The server of claim 13 , wherein:

prior to determining the n-n-number of sequential units, further configured to associate a grammatical type to each word of the digital text; and

wherein the n-number of sequential units are of a predetermined grammatical type.

15. The server of claim 10 , wherein the interrelation analysis comprises one of:

determining a third inclusion parameter of the first context parameter into a third context parameter, wherein the third context parameter is determined by:

further parsing the digital text, by the server, to determine a third lexical unit;

for each entry of the third lexical unit in the text:

selecting, by the server, n-number of sequential units adjacent to the third lexical unit;

generating a third plurality of unit-pairs, the third plurality of unit-pairs comprising the third lexical unit paired with each unit of the n-number of sequential units adjacent to the third lexical unit;

generating, by the server, the third context parameter for the third lexical unit, the third context parameter comprising a frequency of co-occurrence of each unit-pairs of the third plurality of unit-pairs within the one or more sentences of the digital text; and

determining a second similarity parameter between the third context parameter and the second context parameter.

16. The server of claim 15 , wherein:

upon determination that the first inclusion parameter is above the first threshold, the lexical unit relation parameter for the first and second lexical unit is:

a hypernym-hyponym relationship if the inclusion parameter is above a seventh threshold; and

upon determination that the first inclusion parameter and the third inclusion parameter are above a eight threshold, the lexical unit relation parameter for the first, second and third lexical unit is:

a holonym-meronym relationship if the second similarity parameter is below a ninth threshold.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068525/0349 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2017
From: ZELENKOV, YURY GRIGORIEVICH
To: YANDEX LLC
Reel/Frame 042514/0748 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2017
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 042514/0805 →
Priority Claims (1)
RU 2016137530 · Sep 20, 2016 · national
Continuity (1)
Related Publication 20180081874A1 · Mar 22, 2018