IP Library Granted Patent US 11,886,820
Granted Patent B2
US 11,886,820 · App. 17/064,028 · Granted Jan 30, 2024

System and method for machine-learning based extraction of information from documents

Inventors: Sreekanth Menon (Bangalore, IN); Prakash Selvakumar (Bangalore, IN); Sudheesh Sudevan (Thalassery, IN)
Assignee: Genpact Luxembourg S.à r.l. II
G06F40/295G06F18/2148G06F18/2155G06F18/2193G06F18/23213G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,886,820
App. No.
17/064,028
Granted
Jan 30, 2024
Kind
B2
Abstract

A method and system are provided for training a machine-learning (ML) system/module and to provide an ML model. In one embodiment, a method includes using a labeled entities set to train a machine learning (ML) system, to obtain an ML model, and using the trained ML model to predict labels for entities in an unlabeled entities set, yielding a machine-labeled entities set. One or more individual ML models may be trained and used in this way, where each individual ML model corresponds to a respective document source. The document sources can be identified via classification of a corpus of documents. The prediction of labels provides a respective confidence score for each machine-labeled entity. The method also includes selecting from the machine-labeled entities set, a subset of machine-labeled entities having a respective confidence score at least equal to a threshold confidence score; and updating the labeled entities set by adding thereto the selected subset of machine-labeled entities. The method further includes removing from the machine-labeled entities set the selected subset of machine-labeled entities and deleting labels assigned to the entities in the updated machine-labeled entities set to provide the unlabeled entities set for a next iteration. The method also includes, if a termination condition is not reached, repeating the steps above and, otherwise, storing the ML model.

Claims (58)

1. A method for training a machine-learning (ML) system, the method comprising:

(a) providing a seed set of labeled entities as a labeled entities set based on a first cluster of a plurality of clusters of documents and using the labeled entities set to train the ML system, to obtain an ML model;

(b) using the trained ML system to predict labels for entities in an unlabeled entities set, yielding a machine-labeled entities set, the prediction providing a respective confidence score for each machine-labeled entity;

(c) selecting from the machine-labeled entities set, a subset of machine-labeled entities having a respective confidence score at least equal to a threshold confidence score;

(d) updating the labeled entities set by adding thereto the selected subset of machine- labeled entities;

(e) removing from the machine-labeled entities set the selected subset of machine-labeled entities and deleting labels assigned to the entities in the updated machine-labeled entities set to provide the unlabeled entities set for a next iteration;

(f) if a termination condition is not reached, repeating steps (a) through (e), and, otherwise, storing the ML model;

(g) selecting a second cluster from the plurality of clusters; and

(h) repeating the steps (a) through (f) for the second cluster to store a different ML model for the second cluster, wherein providing the seed set in step (a) is based on the second cluster.

2. The method of claim 1 , wherein the seed set comprises manually labeled entities from a subset of documents in the first cluster or a corpus of documents, the subset size not exceeding a specified fraction of the first cluster size or corpus size.

3. The method of claim 2 , wherein the specified fraction is in a range 0.1% up to 10%.

4. The method of claim 2 , further comprising:

selecting the subset of documents using k-means clustering.

5. The method of claim 2 , wherein each cluster of the plurality of clusters is being-associated with a respective document source.

6. The method of claim 1 , further comprising:

using another ML system to:

identify sources of documents in a corpus of documents; and

generate the plurality of clusters based on the identified sources of documents.

7. The method of claim 6 , wherein:

the corpus of documents comprises a plurality of invoices; and

the sources of documents comprise one or more vendors supplying one or more of the plurality of invoices.

8. The method of claim 1 , wherein the termination condition comprises one or more of:

a maximum number of iteration cycles;

a target size of the labeled entities set;

a minimum size of the machine-labeled entities set in one iteration; or

a minimum cumulative size of the machine-labeled entities sets across a plurality of iterations.

9. The method of claim 1 , wherein the ML-system comprises a hybrid of conditional random fields (CRF) and a long short-term memory (LSTM) classifier.

10. The method of claim 1 , wherein the labels that the trained ML system predicts for the entities in the unlabeled entities set are in an inside-outside-beginning (IOB) format.

11. A system for providing a machine-learning (ML) model, the system comprising:

a processor; and

a memory in communication with the processor and comprising instructions which, when executed by the processor, program the processor to:

(a) provide a seed set of labeled entities as a labeled entities set based on a first cluster of a plurality of clusters of documents and use the labeled entities set to train a machine learning (ML) module, to obtain the ML model;

(b) use the trained ML module to predict labels for entities in an unlabeled entities set, yielding a machine-labeled entities set, the prediction providing a respective confidence score for each machine-labeled entity;

(c) select from the machine-labeled entities set, a subset of machine-labeled entities having a respective confidence score at least equal to a threshold confidence score;

(d) update the labeled entities set by adding thereto the selected subset of machine-labeled entities;

(e) remove from the machine-labeled entities set the selected subset of machine-labeled entities and delete labels assigned to the entities in the updated machine-labeled entities set to provide the unlabeled entities set for a next iteration;

(f) if a termination condition is not reached, repeat operations (a) through (e) and, otherwise, store the ML model;

(g) select a second cluster from the plurality of clusters; and

(h) repeat the steps (a) through (f) for the second cluster to store a different ML model for the second cluster, wherein providing the seed set in step (a) is based on the second cluster.

12. The system of claim 11 , wherein the seed set comprises manually labeled entities from a subset of documents in the first cluster or a corpus of documents, the subset size not exceeding a specified fraction of the first cluster size or corpus size.

13. The system of claim 12 , wherein the specified fraction is in a range 0.1% up to 10%.

14. The system of claim 12 , wherein the instructions further program the processor to:

select the subset of documents using k-means clustering.

15. The system of claim 12 , wherein each cluster of the plurality of clusters is associated with a respective document source.

16. The system of claim 11 , wherein the instructions further program the processor to:

use another ML module to:

identify sources of documents in a corpus of documents; and

generate the plurality of clusters based on the identified sources of documents.

17. The system of claim 16 , wherein:

the corpus of documents comprises a plurality of invoices; and

the sources of documents comprise one or more vendors supplying one or more of the plurality of invoices.

18. The system of claim 11 , wherein the termination condition comprises one or more of:

a maximum number of iteration cycles;

a target size of the labeled entities set;

a minimum size of the machine-labeled entities set in one iteration; or

a minimum cumulative size of the machine-labeled entities sets across a plurality of iterations.

19. The system of claim 11 , wherein the ML module comprises a hybrid of conditional random fields (CRF) and a long short-term memory (LSTM) classifier.

20. The system of claim 11 , wherein the labels that the trained ML module predicts for the entities in the unlabeled entities set are in an inside-outside-beginning (IOB) format.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYANCE TYPE OF MERGER PREVIOUSLY RECORDED ON REEL 66511 FRAME 683. ASSIGNOR(S) HEREBY CONFIRMS THE CONVEYANCE TYPE OF ASSIGNMENT. Recorded Feb 26, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 067211/0020 →
MERGER Recorded Feb 7, 2024
From: GENPACT LUXEMBOURG S.À R.L. II
To: GENPACT USA, INC.
Reel/Frame 066511/0683 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2022
From: MENON, SREEKANTH; SELVAKUMAR, PRAKASH; SUDEVAN, SUDHEESH
To: GENPACT LUXEMBOURG S.À R.L. II
Reel/Frame 059572/0033 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2021
From: GENPACT LUXEMBOURG S.À R.L., A LUXEMBOURG PRIVATE LIMITED LIABILITY COMPANY (SOCIÉTÉ À RESPONSABILITÉ LIMITÉE)
To: GENPACT LUXEMBOURG S.À R.L. II, A LUXEMBOURG PRIVATE LIMITED LIABILITY COMPANY (SOCIÉTÉ À RESPONSABILITÉ LIMITÉE)
Reel/Frame 055104/0632 →