IP Library Granted Patent US 9,626,353
Granted Patent B2
US 9,626,353 · App. 14/588,690 · Granted Apr 18, 2017

Arc filtering in a syntactic graph

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,626,353
App. No.
14/588,690
Granted
Apr 18, 2017
Kind
B2
Abstract

The present disclosure provides methods and systems for performing syntactic analysis of a text. In some implementations the method includes performing rough syntactic analysis of the text, generating a graph of generalized constituents of the text and filtering arcs of the graph of generalized constituents with a combination classifier which includes a tree classifier and one or more linear classifiers. The combination classifier is trained using parallel analysis of an untagged two-language text corpus.

Claims (35)

1. A method comprising:

identifying a sentence;

identifying a graph of generalized constituents of the sentence based on rough syntactic analysis of a lexical-morphological structure of the sentence, wherein the graph of generalized constituents comprises arcs and nodes, wherein each of the nodes represents a constituent of the sentence comprising one or more words in the sentence that function as a unit within the sentence, and wherein each of the arcs between a pair of the nodes represents a syntactic slot expressing a type of relationship between lexical values of the pair;

filtering, by a data processing apparatus, the arcs of the graph of generalized constituents using a combination classifier comprising a tree classifier and at least one linear classifier, wherein the tree classifier divides the arcs into clusters based on a predetermined set of symbolic features, and wherein the linear classifier filters the clusters of the arcs based on combinations of numerical features for each of the clusters; and

identifying, by the data processing apparatus, a syntactic structure of the sentence by performing precise syntactic analysis of the sentence based on the graph of generalized constituents of the sentence with the filtered clusters of the arcs.

2. The method of claim 1 , wherein the predetermined set of symbolic features comprises one or more of a type of sentence part of one of the lexical values in the pair, a role of one of the lexical values in the pair, or a number of steps between the pair of the lexical values in a syntactic tree.

3. The method of claim 1 , wherein each of the numerical features comprises a rating for one or more of a lexical meaning, a filler, a punctuation, or a semantic description for the lexical values in the pair.

4. The method of claim 1 , wherein the predetermined set of symbolic features is based on parallel analysis of a two-language text corpus.

5. The method of claim 1 , wherein an order of symbolic features in the predetermined set of symbolic features is determined based on entropy measures of the symbolic features.

6. The method of claim 1 , wherein the tree classifier is based on an Iterative Dichotomiser 3 (ID3) algorithm.

7. The method of claim 1 , wherein the numerical features comprise weights that are based on parallel analysis of a two-language text corpus.

8. Non-transitory computer storage media having instructions stored therein that, when executed by a data processing apparatus, cause the data processing apparatus to:

identify a sentence;

identify a graph of generalized constituents of the sentence based on rough syntactic analysis of a lexical-morphological structure of the sentence, wherein the graph of generalized constituents comprises arcs and nodes, wherein each of the nodes represents a constituent of the sentence comprising one or more words in the sentence that function as a unit within the sentence, and wherein each of the arcs between a pair of the nodes represents a syntactic slot expressing a type of relationship between lexical values of the pair;

filter, by the data processing apparatus, the arcs of the graph of generalized constituents using a combination classifier comprising a tree classifier and at least one linear classifier, wherein the tree classifier divides the arcs into clusters based on a predetermined set of symbolic features, and wherein the linear classifier filters the clusters of the arcs based on combinations of numerical features for each of the clusters; and

identify, by the data processing apparatus, a syntactic structure of the sentence by performing precise syntactic analysis of the sentence based on the graph of generalized constituents of the sentence with the filtered clusters of the arcs.

9. The non-transitory computer storage media of claim 8 , wherein the predetermined set of symbolic features comprises one or more of a type of sentence part of one of the lexical values in the pair, a role of one of the lexical values in the pair, or a number of steps between the pair of the lexical values in a syntactic tree.

10. The non-transitory computer storage media of claim 8 , wherein each of the numerical features comprises a rating for one or more of a lexical meaning, a filler, a punctuation, or a semantic description for the lexical values in the pair.

11. The non-transitory computer storage media of claim 8 , wherein the predetermined set of symbolic features is based on parallel analysis of a two-language text corpus.

12. The non-transitory computer storage media of claim 8 , wherein an order of symbolic features in the predetermined set of symbolic features is determined based on entropy measures of the symbolic features.

13. The non-transitory computer storage media of claim 8 , wherein the tree classifier is based on an Iterative Dichotomiser 3 (ID3) algorithm.

14. The non-transitory computer storage media of claim 8 , wherein the numerical features comprise weights that are based on parallel analysis of a two-language text corpus.

15. A system comprising:

a data processing apparatus; and

a computer-readable medium having instructions stored therein that, when executed by the data processing apparatus, cause the data processing apparatus to:

identify a sentence;

identify a graph of generalized constituents of the sentence based on rough syntactic analysis of a lexical-morphological structure of the sentence, wherein the graph of generalized constituents comprises arcs and nodes, wherein each of the nodes represents a constituent of the sentence comprising one or more words in the sentence that function as a unit within the sentence, and wherein each of the arcs between a pair of the nodes represents a syntactic slot expressing a type of relationship between lexical values of the pair;

filter the arcs of the graph of generalized constituents using a combination classifier comprising a tree classifier and at least one linear classifier, wherein the tree classifier divides the arcs into clusters based on a predetermined set of symbolic features, and wherein the linear classifier filters the clusters of the arcs based on combinations of numerical features for each of the clusters; and

identify a syntactic structure of the sentence by performing precise syntactic analysis of the sentence based on the graph of generalized constituents of the sentence with the filtered clusters of the arcs.

16. The system of claim 15 , wherein the predetermined set of symbolic features comprises one or more of a type of sentence part of one of the lexical values in the pair, a role of one of the lexical values in the pair, or a number of steps between the pair of the lexical values in a syntactic tree.

17. The system of claim 15 , wherein each of the numerical features comprises a rating for one or more of a lexical meaning, a filler, a punctuation, or a semantic description for the lexical values in the pair.

18. The system of claim 15 , wherein the predetermined set of symbolic features is based on parallel analysis of a two-language text corpus.

19. The system of claim 15 , wherein an order of symbolic features in the predetermined set of symbolic features is determined based on entropy measures of the symbolic features.

20. The system of claim 15 , wherein the tree classifier is based on an Iterative Dichotomiser 3 (ID3) algorithm.

21. The system of claim 15 , wherein the numerical features comprise weights that are based on parallel analysis of a two-language text corpus.

Assignments (5)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR DOC. DATE PREVIOUSLY RECORDED AT REEL: 042706 FRAME: 0279. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 25, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 043676/0232 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2017
From: ABBYY INFOPOISK LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 042706/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2015
From: ANISIMOVICH, KONSTANTIN; ZUEV, KONSTANTIN ALEKSEEVICH
To: ABBYY INFOPOISK LLC
Reel/Frame 034755/0321 →