IP Library › Granted Patent US 12,205,395
Granted Patent B1
US 12,205,395 · App. 17/405,127 · Granted Jan 21, 2025

Key-value extraction from documents

Inventors: Adrian Yunpfei Lam (South San Francisco, CA); Chiao-Lun Cheng (San Francisco, CA); Alexandre Matton (San Francisco, CA)
Assignee: Scale AI, Inc.
G06V30/416G06N20/00G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,395
App. No.
17/405,127
Granted
Jan 21, 2025
Kind
B1
Abstract

One embodiment of the present invention sets forth a technique for extracting data from a document. The technique includes determining a first set of features associated with the document, wherein the first set of features comprises a set of region proposals that bound one or more portions of text within the document. The technique also includes applying a first machine learning model to the first set of features to generate a set of predictions associated with one or more key-value pairs included in the document. The technique further includes extracting the one or more key-value pairs from the document based on the set of predictions.

Claims (43)

1. A computer-implemented method for extracting data from a document, the method comprising:

determining a first set of features associated with the document, wherein the first set of features comprises a set of region proposals that bound one or more portions of text within the document;

applying a first machine learning model to the first set of features to generate a set of predictions comprising a set of scores associated with one or more key-value pairs, wherein the set of scores includes, for each region proposal in the set of region proposals, a first score that represents a probability that the region proposal includes text associated with a single key in the one or more key-value pair, a second score that represents a probability that the region proposal includes text associated with a single value in the one or more key-value pairs, and a third score that represents a probability that the region proposal includes text that is unrepresentative of any single key or any single value in the one or more key-value pairs; and

extracting the one or more key-value pairs from the document based on the set of predictions.

2. The computer-implemented method of claim 1 , wherein extracting the one or more key-value pairs from the document based on the set of predictions comprises extracting the text from a region proposal associated with a score that meets a threshold.

3. The computer-implemented method of claim 1 , further comprising:

generating a set of synthetic key-value pairs based on one or more key-value samples extracted from one or more documents; and

inputting the set of synthetic key-value pairs as training data for the first machine learning model.

4. The computer-implemented method of claim 3 , wherein generating the set of synthetic key-value pairs comprises:

training a generative model based on the one or more key-value samples; and

executing the trained generative model to produce the set of synthetic key-value pairs.

5. The computer-implemented method of claim 3 , wherein generating the set of synthetic key-value pairs comprises generating a key-value pair that includes a formatting associated with the one or more key-value samples and text that differs from the one or more key-value samples.

6. The computer-implemented method of claim 1 , wherein extracting the one or more key-value pairs from the document based on the set of predictions comprises:

generating a second set of features based on the set of predictions;

applying a second machine learning model to a second set of features to generate one or more bounding boxes for one or more words within the document; and

extracting the one or more key-value pairs based on the one or more bounding boxes.

7. The computer-implemented method of claim 6 , wherein each of the one or more bounding boxes includes at least one of a value or a key-value pair.

8. The computer-implemented method of claim 6 , wherein the second set of features comprises at least one of the document or the set of predictions.

9. The computer-implemented method of claim 1 , wherein each region proposal included in the set of region proposals represents a grouping of one or more bounding boxes for the one or more portions of text within the document.

10. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

determining a first set of features associated with a document, wherein the first set of features comprises a set of region proposals that bound one or more portions of text within the document;

applying a first machine learning model to the first set of features to generate a set of predictions comprising a set of scores associated with one or more key-value pairs, wherein the set of scores includes, for each region proposal in the set of region proposals, a first score that represents a probability that the region proposal includes text associated with a single key in the one or more key-value pair, a second score that represents a probability that the region proposal includes text associated with a single value in the one or more key-value pairs, and a third score that represents a probability that the region proposal includes text that is unrepresentative of any single key or any single value in the one or more key-value pairs; and

extracting the one or more key-value pairs from the document based on the set of predictions.

11. The one or more non-transitory computer readable media of claim 10 , wherein extracting the one or more key-value pairs from the document based on the set of predictions comprises extracting the text from a region proposal associated with a score that meets a threshold.

12. The one or more non-transitory computer readable media of claim 10 , wherein the operations further comprise filtering the set of region proposals based on an overlap in two or more region proposals prior to extracting the one or more key-value pairs from the document.

13. The one or more non-transitory computer readable media of claim 10 , wherein the operations further comprise:

generating a set of synthetic key-value pairs based on one or more key-value samples extracted from one or more documents; and

inputting the set of synthetic key-value pairs as training data for the first machine learning model.

14. The one or more non-transitory computer readable media of claim 13 , wherein generating the set of synthetic key-value pairs comprises:

training a generative model based on the one or more key-value samples; and

executing the trained generative model to produce the set of synthetic key-value pairs.

15. The one or more non-transitory computer readable media of claim 13 , wherein generating the set of synthetic key-value pairs comprises generating a key-value pair that includes a formatting associated with the one or more key-value samples and text that differs from the one or more key-value samples.

16. The one or more non-transitory computer readable media of claim 10 , wherein extracting the one or more key-value pairs from the document based on the set of predictions comprises:

generating a second set of features based on the set of predictions;

applying a second machine learning model to the second set of features to generate one or more bounding boxes for one or more words within the document; and

extracting the one or more key-value pairs based on the one or more bounding boxes.

17. The one or more non-transitory computer readable media of claim 10 , wherein the first set of features further comprises a set of feature maps associated with the set of region proposals.

18. A system, comprising:

a memory that stores instructions, and

a processor that is coupled to the memory and, when executing the instructions, is configured to:

determine a first set of features associated with a document, wherein the first set of features comprises a set of region proposals that bound one or more portions of text within the document;

apply a first machine learning model to the first set of features to generate a set of predictions comprising a set of scores associated with one or more key-value pairs, wherein the set of scores includes, for each region proposal in the set of region proposals, a first score that represents a probability that the region proposal includes text associated with a single key in the one or more key-value pair, a second score that represents a probability that the region proposal includes text associated with a single value in the one or more key-value pairs, and a third score that represents a probability that the region proposal includes text that is unrepresentative of any single key or any single value in the one or more key-value pairs; and

extract the one or more key-value pairs from the document based on the set of predictions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2021
From: LAM, ADRIAN YUNPFEI; CHENG, CHIAO-LUN; MATTON, ALEXANDRE
To: SCALE AI, INC.
Reel/Frame 057255/0534 →
References Cited (16)
US 10445569B1 · Lin et al. · 2019 [cited by applicant]
US 10839245B1 · Dhillon et al. · 2020 [cited by applicant]
US 10896357B1 · Corcoran · 2021 [cited by examiner]
US 10915788B2 · Hoehne et al. · 2021 [cited by applicant]
US 11182604B1 · Methaniya et al. · 2021 [cited by applicant]
US 11810380B2 · Arroyo et al. · 2023 [cited by applicant]
US 20200159820A1 · Rodriguez et al. · 2020 [cited by applicant]
US 20200160050A1 · Bhotika et al. · 2020 [cited by applicant]
US 20210073533A1 · Ast · 2021 [cited by examiner]
US 20210383106A1 · Maggio · 2021 [cited by examiner]
US 20220036063A1 · Bhuyan · 2022 [cited by examiner]
US 20220300834A1 · Zeng et al. · 2022 [cited by applicant]
Liao et al., “Real-time Scene Text Detection with Differentiable Binarization”, trainarXiv: 1911.08947v2 [cs.CV] Dec. 3, 2019, 8 pages. [cited by applicant]
Amazon Textract Developer Guide, https://docs.aws.amazon.com/textract/latest/dg/textract-dg.pdf, 240 pages. [cited by applicant]
Non-final Office Action received for U.S. Appl. No. 17/405,141 dated May 9, 2024, 21 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/405,141 dated Aug. 29, 2024, 12 pages. [cited by applicant]
Cited By (3)
US 12,361,736 US 12,572,740 US 12,705,916