IP Library › Granted Patent US 12,158,900
Granted Patent B2
US 12,158,900 · App. 17/976,439 · Granted Dec 3, 2024

Extracting information from documents using automatic markup based on historical data

Inventor: Stanislav Semenov (Moscow, RU)
Assignee: ABBYY Development Inc.
G06F16/288G06F16/215G06F16/24578G06F16/248G06F16/285G06F16/335G06F16/38G06F16/94G06N3/08G06N3/082G06N5/046G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,158,900
App. No.
17/976,439
Granted
Dec 3, 2024
Kind
B2
Abstract

Mechanisms for document processing and analysis can include receiving a document and identifying, in a data structure, a record corresponding to the document. The record can include one or more entries, where each entry contains data reflecting a respective item of information extracted from a corresponding part of the document. The mechanisms can include determining for each entry of the record, a corresponding degree of association between the entry and a respective item of information referenced by the entry. They can further include updating the corresponding degrees of association, and selecting, among the corresponding degrees of association, a set of corresponding degrees of association whose aggregate degree of association satisfies a criterion.

Claims (36)

1. A method comprising:

receiving, by a processing device, a document;

identifying, in a data structure, a record corresponding to the document, the record comprising one or more entries, each entry containing data reflecting a respective item of information extracted from a corresponding part of the document;

for each entry of the record, determining a first embedding of a first character string contained by the entry;

identifying, among a set of embeddings corresponding to character strings of the document, a second embedding that is closest in a vector space to the first embedding;

identifying, in the document, a second character string encoded by the second embedding;

determining a degree of association between the entry and the second character string;

selecting, among a plurality of degrees of association between entries of the data structure and corresponding character strings, a set of degrees of association whose aggregate degree of association satisfies a criterion; and

training, using the set of degrees of association, a machine learning model to extract information from new documents.

2. The method of claim 1 , further comprising:

updating the degree of association by performing at least one of:

identifying character strings using a word search, detecting fields with a neural network model, or receiving identification of character strings from a user interface.

3. The method of claim 2 , wherein the word search comprises: searching a body of text in the document for a character string from the entry of the record by performing one or more of an exact string search, a prefix search, an approximate string search, or an approximate prefix search.

4. The method of claim 2 , wherein detecting the fields with the neural network model comprises: marking one or more areas in a document as fields, each field associated with a class of information, and matching a character string from the entry in the record with one or more fields based on the class of information associated with each field.

5. The method of claim 1 , further comprising performing at least one of:

eliminating degrees of association that are lower than a predefined threshold value or combining degrees of association resulting from different character string identification methods.

6. The method of claim 1 , wherein selecting the set of degrees of association that satisfies a criterion comprises: applying a bijective function to a set of entries from the record and a set of items of information identified in the document.

7. A system comprising:

a memory device comprising a data structure;

a processor coupled to the memory device, the processor configured to perform operations comprising:

receiving a document;

identifying, in a data structure, a record corresponding to the document, the record comprising one or more entries, each entry containing data reflecting a respective item of information extracted from a corresponding part of the document;

for each entry of the record, determining a first embedding of a first character string contained by the entry;

identifying, among a set of embeddings corresponding to character strings of the document, a second embedding that is closest in a vector space to the first embedding;

identifying, in the document, a second character string encoded by the second embedding;

determining a degree of association between the entry and the second character string;

selecting, among a plurality of degrees of association between entries of the data structure and corresponding character strings, a set of degrees of association whose aggregate degree of association satisfies a criterion; and

training, using the set of degrees of association, a machine learning model to extract information from new documents.

8. The system of claim 7 , wherein the operations further comprise:

updating the degree of association by performing at least one of:

identifying character strings using a word search, detecting fields with a neural network model, or receiving identification of character strings from a user interface.

9. The system of claim 8 , wherein the word search comprises: searching a body of text in the document for a character string from the entry from the record by performing one or more of an exact string search, a prefix search, an approximate string search, or an approximate prefix search.

10. The system of claim 8 , wherein detecting the fields with the neural network model comprises: marking one or more areas in a document as fields, each field associated with a class of information, and matching a character string from the entry in the record with one or more fields based on the class of information associated with each field.

11. The system of claim 7 , wherein the operations further comprise:

performing at least one of: eliminating corresponding degrees of association that are lower than a predefined threshold value or combining degrees of association resulting from different character string identification methods.

12. The system of claim 7 , wherein selecting the set of degrees of association that satisfies a criterion comprises: applying a bijective function to a set of entries from the record and a set of items of information identified in the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: SEMENOV, STANISLAV
To: ABBYY DEVELOPMENT INC.
Reel/Frame 061828/0874 →
Continuity (1)
Related Publication 20240143632A1 · May 2, 2024
Cited By (1)
US 1,089,276