IP Library Granted Patent US 10,445,428
Granted Patent B2
US 10,445,428 · App. 15/852,765 · Granted Oct 15, 2019

Information object extraction using combination of classifiers

Inventors: Stepan Evgenyevich Matskevich (Korolev, RU); Dmitry Andreevich Sukhodolov (Dolgoprudniy, RU); Anatoly Sergeevich Starostin (Moscow, RU)
Assignee: ABBYY Production LLC
G06F17/2785G06F17/271G06F17/241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,445,428
App. No.
15/852,765
Granted
Oct 15, 2019
Kind
B2
Abstract

Systems and methods for information extraction from natural language texts using a combination of classifier models. An example method may comprise: producing, by performing syntactico-semantic analysis of a natural language text, a plurality of syntactico-semantic structures representing the natural language text; identifying, using a first classifier model to process a first plurality of classification attributes derived from the syntactico-semantic structures, a plurality of core constituents, such that each core constituent of the plurality of core constituents is associated with a span of a plurality of spans, wherein each span represents an attribute of an information object of a specified ontology class; identifying, using a second classifier model to process a second plurality of classification attributes derived from the syntactico-semantic structures, child constituents of each of the plurality of core constituents; and determining, using a third classifier model to process a third plurality of classification attributes derived from the syntactico-semantic structures, whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.

Claims (45)

1. A method, comprising:

performing, by a computer system, syntactico-semantic analysis of a natural language text;

processing, by the computer system, based on a first classifier model, a first plurality of classification attributes derived from a plurality of syntactico-semantic structures produced by the syntactico-semantic analysis, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;

processing, by the computer system, based on a second classifier model, a second plurality of classification attributes derived from the syntactico-semantic structures, wherein the second classifier model identifies a plurality of child constituents of each of the plurality of core constituents;

identifying a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes the identified child constituents of the core constituent; and

processing, by the computer system, based on a third classifier model, a third plurality of classification attributes derived from the syntactico-semantic structures, wherein the third classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.

2. The method of claim 1 , further comprising:

utilizing, for performing a natural language processing task, the information object attributes associated with the first span and the second span.

3. The method of claim 1 , further comprising:

displaying, in visual association with a first projection of the first span and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and

accepting user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.

4. The method of claim 1 , wherein each syntacitco-semantic structure of the plurality of syntacitco-semantic structures is represented by a graph comprising a plurality of nodes corresponding to a plurality of semantic classes and a plurality of edges corresponding to a plurality of semantic relationships.

5. The method of claim 1 , further comprising:

determining, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.

6. The method of claim 1 , wherein the first classifier model yields a value representing a likelihood of a candidate node representing a core constituent of a span that represents an attribute of an information object of the specified ontology class.

7. The method of claim 1 , wherein the first plurality of classification attributes include attributes of a candidate core constituent and at least one of: a parent node of the candidate core constituent, a child node of the candidate core constituent, or a sibling node of the candidate core constituent.

8. The method of claim 1 , wherein the second classifier model yields a value representing a likelihood of a candidate child constituent belonging to a span associated with a specified core constituent.

9. The method of claim 1 , wherein the second plurality of classification attributes include attributes of a candidate child constituent and at least one of: a parent node of the candidate child constituent, a child node of the candidate child constituent, or a sibling node of the candidate child constituent.

10. The method of claim 1 , wherein the third classifier model yields a value representing a likelihood of the first span and the second span being associated with the same information object.

11. The method of claim 1 , wherein the third plurality of classification attributes include attributes of nodes of the first span and attributes of nodes of the second span.

12. The method of claim 1 , further comprising:

producing the first plurality of classification attributes by traversing, according to a pre-defined traversal path, at least one syntactico-semantic structure of the plurality of syntactico-semantic structures, wherein the traversal path specifies a plurality of nodes whose attributes are to be included into the first plurality of classification attributes.

13. A system, comprising:

a memory;

a processor, coupled to the memory, the processor configured to:

perform syntactico-semantic analysis of a natural language text;

process, based on a first classifier model, a first plurality of classification attributes derived from a plurality of syntactico-semantic structures produced by the syntactico-semantic analysis, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;

process, based on a second classifier model, a second plurality of classification attributes derived from the syntactico-semantic structures, wherein the second classifier model identifies a plurality of child constituents of each of the plurality of core constituents;

identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes the identified child constituents of the core constituent; and

process, based on a third classifier model, a third plurality of classification attributes derived from the syntactico-semantic structures, wherein the third classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.

14. The system of claim 13 , wherein the processor is further configured to:

utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.

15. The system of claim 13 , wherein the first classifier model yields a value representing a likelihood of a candidate node representing a core constituent of a span that represents an attribute of an information object of the specified ontology class.

16. The system of claim 13 , wherein the second classifier model yields a value representing a likelihood of a candidate child constituent belonging to a span associated with a specified core constituent.

17. The system of claim 13 , wherein the third classifier model yields a value representing a likelihood of the first span and the second span being associated with the same information object.

18. A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:

perform syntactico-semantic analysis of a natural language text;

process, based on a first classifier model, a first plurality of classification attributes derived from a plurality of syntactico-semantic structures produced by the syntactico-semantic analysis, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;

process, based on a second classifier model, a second plurality of classification attributes derived from the syntactico-semantic structures, wherein the second classifier model identifies a plurality of child constituents of each of the plurality of core constituents;

identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes the identified child constituents of the core constituent; and

process, based on a third classifier model, a third plurality of classification attributes derived from the syntactico-semantic structures, wherein the third classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.

19. The computer-readable non-transitory storage medium of claim 18 , further comprising executable instructions to cause the computer system to:

utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.

20. The computer-readable non-transitory storage medium of claim 18 , further comprising executable instructions to cause the computer system to:

determine, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
MERGER Recorded Jan 24, 2019
From: ABBYY DEVELOPMENT LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 048129/0558 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2018
From: MATSKEVICH, STEPAN EVGENYEVICH; SUKHODOLOV, DMITRY ANDREEVICH; STAROSTIN, ANATOLY SERGEEVICH
To: ABBYY DEVELOPMENT LLC
Reel/Frame 044542/0549 →
Priority Claims (1)
RU 2017143154 · Dec 11, 2017 · national
Continuity (1)
Related Publication 20190179897A1 · Jun 13, 2019