IP Library Granted Patent US 8,812,296
Granted Patent B2
US 8,812,296 · App. 11/769,478 · Granted Aug 19, 2014

Method and system for natural language dictionary generation

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,812,296
App. No.
11/769,478
Granted
Aug 19, 2014
Kind
B2
Abstract

A method and computer system for analyzing a text corpus in a natural language is provided. An initial morphological description having word inflection rules for various groups of words in the natural language is created by a linguist. A plurality of text corpuses are analyzed to obtain information on the occurrence of a plurality of word forms for each word token in each text corpus. A morphological dictionary which contains information about each base form and word inflection rules for each word token with verified hypothesis is generated.

Claims (76)

1. A method for a computer system to create a morphological dictionary for a natural language, the method comprising:

identifying a word token in a text corpus;

applying by the computer system one or more paradigm rules to the word token;

generating by the computer system one or more hypotheses about a part of speech for a base form of the word token;

searching by the computer system for one or more word inflected forms corresponding to the base form of the word token;

verifying by the computer system a hypothesis of the one or more hypotheses for the base form of the word token;

adding by the computer system at least one grammatical value and at least one inflection paradigm to the base form of the word token based at least in part on the verified hypothesis;

obtaining by the computer system one or more morphological descriptions for the word token based at least in part on the verified hypothesis; and

adding the base form of the word token with the one or more morphological descriptions to the morphological dictionary of the natural language.

2. The method of claim 1 , further comprising creating an initial morphological description having word inflection rules for groups of words in the natural language.

3. The method of claim 1 , wherein each of the one or more morphological descriptions comprises one or more word inflection rules.

4. The method of claim 1 , wherein the one or more morphological descriptions comprise one or more word formation rules.

5. The method of claim 1 , wherein the one or more morphological descriptions comprise a grammatical system of the natural language.

6. The method of claim 5 , wherein the grammatical system of the natural language comprises a set of grammatical categories and grammemes associated with the set of grammatical categories.

7. The method of claim 1 , further comprising:

using additional ratings for the hypothesis when an initial verification of the hypothesis is unsuccessful.

8. The method of claim 7 , wherein the additional ratings are obtained by checking the base form against a word list of a dictionary.

9. The method of claim 7 , wherein the additional ratings are obtained by checking the base form with a spelling checker system.

10. The method of claim 1 , wherein the one or more paradigm rules applied to the word token include one or more rules for changing each word token, and wherein the one or more rules include adding an affix, removing an affix, adding an ending, removing an ending, or combinations thereof.

11. The method of claim 1 , wherein the at least one grammatical value added to the base form of the word token includes information about the part of speech for the word token.

12. The method of claim 1 , wherein the at least one inflection paradigms added to the base form of the word token comprises inflection rules for the word token.

13. The method of claim 1 , wherein the natural language comprises English, French, German, Italian, Russian, Spanish, Ukrainian, Dutch, Danish, Swedish, Finnish, Portuguese, Slovak, Polish, Czech, Hungarian, Lithuanian, Latvian, Estonian, Greek, Bulgarian, Turkish, Tatar, Hindi, Serbian, Croatian, Romanian, Slovenian, Macedonian, Japanese, Korean, Arabic, Hebrew, or Swahili.

14. The method of claim 1 , wherein verifying the hypothesis for the base form of the word token further comprises determining which of the one or more word inflected forms corresponding to the base form of the word token occur in the text corpus.

15. The method of claim 14 , wherein the hypothesis is considered to be verified when a specified threshold portion of the one or more word inflected forms occurs in the text corpus.

16. The method of claim 15 , wherein the specified threshold portion is specified as all of the other word inflected forms.

17. The method of claim 1 , wherein the method further comprises:

generating ratings for each of the one or more generated hypotheses, wherein verifying the hypothesis of the one or more hypotheses for the base form of the word token is based on the generated ratings.

18. A method for a computer system to generate a morphological dictionary for a natural language, the method comprising:

creating by the computer system an initial morphological description having word inflection rules for groups of words in the natural language;

analyzing by the computer system a plurality of text corpuses in the natural language, wherein the analyzing includes:

identifying a word token in the plurality of text corpuses;

applying one or more paradigm rules to the word token;

generating one or more hypotheses about one or more parts of speech of a base form of the word token;

searching for one or more word inflected forms corresponding to the base form of the word token;

verifying a hypothesis of the one or more hypotheses for the base form of the word token based on ratings;

adding at least one grammatical value and at least one inflection paradigm to the base form of the word token based at least in part on the verified hypothesis; and

obtaining one or more morphological descriptions for the word token with a verified hypothesis; and

adding the base form of the word token with the one or more morphological descriptions to the morphological dictionary.

19. The method of claim 18 , wherein the one or more morphological descriptions includes one or more word inflection rules.

20. The method of claim 18 , wherein the one or more morphological descriptions includes one or more word formation rules.

21. The method of claim 18 , wherein the one or more morphological descriptions includes a grammatical system of the natural language.

22. The method of claim 18 , wherein the natural language comprises English, French, German, Italian, Russian, Spanish, Ukrainian, Dutch, Danish, Swedish, Finnish, Portuguese, Slovak, Polish, Czech, Hungarian, Lithuanian, Latvian, Estonian, Greek, Bulgarian, Turkish, Tatar, Hindi, Serbian, Croatian, Romanian, Slovenian, Macedonian, Japanese, Korean, Arabic, Hebrew, or Swahili.

23. A computer readable non-transitory medium comprising instructions for causing a computing system to carry out operations for analyzing a text corpus in a natural language, the operations comprising:

identifying a word token in the text corpus;

applying one or more paradigm rules to the word token;

generating one or more hypotheses about a part of speech of a base form of the word token;

searching for one or more word inflected forms corresponding to the base form of the word token;

verifying a hypothesis of the one or more hypotheses for the base form of the word token;

adding at least one grammatical value and at least one inflection paradigm to the base form of the word token based at least in part on the verified hypothesis; and

obtaining one or more morphological descriptions for the word token based at least in part on the verified hypothesis.

24. The computer readable non-transitory medium of claim 23 , wherein the natural language comprises English, French, German, Italian, Russian, Spanish, Ukrainian, Dutch, Danish, Swedish, Finnish, Portuguese, Slovak, Polish, Czech, Hungarian, Lithuanian, Latvian, Estonian, Greek, Bulgarian, Turkish, Tatar, Hindi, Serbian, Croatian, Romanian, Slovenian, Macedonian, Japanese, Korean, Arabic, Hebrew, or Swahili.

25. A computer readable non-transitory medium comprising instructions for causing a computing system to carry out operations to generate a morphological dictionary for a natural language, the instructions comprising:

analyzing a plurality of text corpuses in the natural language, wherein the analyzing includes:

identifying a word token in the plurality of text corpuses;

applying one or more paradigm rules to the word token;

generating one or more hypotheses based in part on a part of speech of a base form of the word token;

searching for one or more word inflected forms corresponding to the base form of the word token;

verifying a hypothesis of the one or more hypotheses for the base form of the word token based on ratings;

adding at least one grammatical value and at least one inflection paradigm to the base form of the word token; and

obtaining one or more morphological descriptions for the word token based at least in part on the verified hypothesis; and

adding the base form of the word token with the one or more morphological descriptions to the morphological dictionary.

26. The computer readable non-transitory medium of claim 25 , wherein the natural language comprises English, French, German, Italian, Russian, Spanish, Ukrainian, Dutch, Danish, Swedish, Finnish, Portuguese, Slovak, Polish, Czech, Hungarian, Lithuanian, Latvian, Estonian, Greek, Bulgarian, Turkish, Tatar, Hindi, Serbian, Croatian, Romanian, Slovenian, Macedonian, Japanese, Korean, Arabic, Hebrew, or Swahili.

27. A system for capturing data from a document image, the system comprising:

a processor; and

a memory coupled to the processor and in electronic communication with the imaging component, the memory configured with instructions for causing the processor to:

identify a word token in a text corpus;

apply by the computer system one or more paradigm rules to the word token;

generate by the computer system one or more hypotheses about a part of speech for a base form of the word token;

search by the computer system for one or more word inflected forms corresponding to the base form of the word token;

verify by the computer system a hypothesis of the one or more hypotheses for the base form of the word token;

add by the computer system at least one grammatical value and at least one inflection paradigm to the base form of the word token based at least in part on the verified hypothesis;

obtain by the computer system one or more morphological descriptions for the word token based at least in part on the verified hypothesis; and

add the base form of the word token with the one or more morphological descriptions to the morphological dictionary.

28. The system of claim 27 , which is operable to create an initial morphological description having word inflection rules for groups of words.

29. The system of claim 27 , wherein each of the one or more morphological descriptions comprises one or more word inflection rules.

30. The system of claim 27 , wherein the one or more morphological descriptions comprise one or more word formation rules.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR DOC. DATE PREVIOUSLY RECORDED AT REEL: 042706 FRAME: 0279. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 25, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 043676/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 042706/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2013
From: ABBYY SOFTWARE LTD.
To: ABBYY INFOPOISK LLC
Reel/Frame 031082/0677 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME AND ADDRESS PREVIOUSLY RECORDED ON REEL 019488 FRAME 0655. ASSIGNOR(S) HEREBY CONFIRMS THE CORRECT NAME AND ADDRESS ARE AS FOLLOWS: ABBYY SOFTWARE LTD. STASIKRATOUS 29, OFFICE 202 CY 1065, NICOSIA, CYPRUS. Recorded Jul 27, 2007
From: SELEGEY, VLADIMIR; MARAMCHIN, ALEXEY
To: ABBYY SOFTWARE LTD.
Reel/Frame 019604/0974 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2007
From: SELEGEY, VLADIMIR; MARAMCHIN, ALEXEY
To: AABBYY SOFTWARE LTD.
Reel/Frame 019488/0665 →