IP Library Granted Patent US 11,798,301
Granted Patent B1
US 11,798,301 · App. 18/047,008 · Granted Oct 24, 2023

Compositional pipeline for generating synthetic training data for machine learning models to extract line items from OCR text

Inventor: Tharathorn Rimchala (Mountain View, CA)
Assignee: INTUIT INC.
G06V30/19147G06N20/00G06V30/30G06V30/414G06V30/416G06V30/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,798,301
App. No.
18/047,008
Granted
Oct 24, 2023
Kind
B1
Abstract

Systems and methods of generating synthetic training data for machine learning models. First, line items in source documents such as bills, invoices, and or receipts are identified and labeled. The identification and labeling generate labeled documents. Then, in the labeled documents, the line items are augmented by adding, deleting, and or swapping line items to generate synthetic training documents. An addition operation randomly selects one or more line items and adds the selected line item(s) to the same labeled document or another labeled document. A deletion operation randomly deletes one or more line items. A swapping operation randomly swaps line items in a single labeled document or across different labeled documents. These operations can generate synthetic labeled documents of any length, which form synthetic training data for training the machine learning models.

Claims (68)

1. A method to generate training data for machine learning models, the method performed by a processor and comprising:

detecting a line items block in a source document;

labeling individual line items in the line items block to generate a labeled document from the source document; and

performing a line-wise augmentation of the individual line items by an adding operation to generate a labeled synthetic document comprising an augmented line items block from the line items block of the labeled document, wherein both the labeled document and the labeled synthetic document are configured to be used as the training data, the adding operation comprising:

randomly selecting a line item in the line items block of the labeled document; and

adding the randomly selected line item to the line items block to generate the labeled synthetic document with the augmented line items block.

2. The method of claim 1 , wherein the adding operation further comprises:

randomly selecting a second line item in a second line items block of a second labeled document different from the labeled document; and

adding the randomly selected second line item to the line items block to generate the labeled synthetic document with the augmented line items block.

3. The method of claim 1 , the line-wise augmentation of the individual line items being further performed by a deleting operation comprising:

randomly selecting a second line item in the line items block of the labeled document; and

deleting the randomly selected second line item from the line items block to generate the labeled synthetic document with the augmented line items block.

4. The method of claim 1 , the line-wise augmentation of the individual line items being further performed by a swapping operation comprising:

randomly selecting a second line item in the line items block of the labeled document;

randomly selecting a third line item in the line items block of the labeled document; and

swapping positions of the third line item and the second line item in the line items block to generate the labeled synthetic document with the augmented line items block.

5. The method of claim 1 , the line-wise augmentation of the individual line items being further performed by a swapping operation comprising:

randomly selecting a second line item in the line items block of the labeled document;

randomly selecting a third line item in a second line items block of a second document; and

swapping positions of the second line item and the third line item in the line items block and the second line items block to generate the labeled synthetic document with the augmented line items block and to generate a second synthetic labeled document from the second document.

6. The method of claim 1 , further comprising:

recalculating a field value outside of the items block in response to performing the adding operation.

7. The method of claim 6 , wherein the labeled document comprises at least one of a bill, invoice, or receipt, and wherein recalculating the field value comprises:

recalculating at least one of a subtotal field, a tax field, or a total field.

8. The method of claim 1 , wherein detecting the line items block comprises:

heuristically determining geometric bounds of the line items block based on other labeled information blocks; or

using a pre-trained table detection machine learning model.

9. The method of claim 1 , wherein labeling the individual line items comprises:

extracting text in the line items block using optical character recognition;

identifying numeric strings from the extracted text;

determining that the identified numeric strings satisfy arithmetic constraints;

using vertical positions of the numeric strings to define the individual line items; and

labeling the numeric strings as line item amounts and corresponding text as line item description.

10. A system comprising:

a non-transitory storage medium storing computer program instructions; and

one or more processors configured to execute the computer program instructions to cause operations comprising:

detecting line items block in a source document;

labeling individual line items in the line items block to generate a labeled document from the source document; and

performing a line-wise augmentation of the individual line items by an adding operation to generate a labeled synthetic document comprising an augmented line items block from the line items block of the labeled document, wherein both the labeled document and the labeled synthetic document are configured to be used as training data, the adding operation comprising:

randomly selecting a line item in the line items block of the labeled document; and

adding the randomly selected line item to the line items block to generate the labeled synthetic document with the augmented line items block.

11. The system of claim 10 , wherein the adding operation further comprises:

randomly selecting a second line item in a second line items block of a second labeled document different from the labeled document; and

adding the randomly selected second line item to the line items block to generate the labeled synthetic document with the augmented line items block.

12. The system of claim 10 , the line-wise augmentation of the individual line items being further performed by a deleting operation comprising:

randomly selecting a second line item in the line items block of the labeled document; and

deleting the randomly selected second line item from the line items block to generate the labeled synthetic document with the augmented line items block.

13. The system of claim 10 , the line-wise augmentation of the individual line items being further performed by a swapping operation comprising:

randomly selecting a second line item in the line items block of the labeled document;

randomly selecting a third line item in the line items block of the labeled document; and

swapping positions of the third line item and the second line item in the line items block to generate the labeled synthetic document with the augmented line items block.

14. The system of claim 10 , the line-wise augmentation of the individual line items being further performed by a swapping operation comprising:

randomly selecting a second line item in the line items block of the labeled document;

randomly selecting a third line item in second line items block of a second document; and

swapping positions of the second line item and the third line item in the line items block and the second line items block to generate the labeled synthetic document with the augmented line items block and to generate a second synthetic labeled document from the second document.

15. The system of claim 10 , further comprising:

recalculating a field value outside of the items block in response to performing the adding operation.

16. The system of claim 15 , wherein the labeled document comprises at least one of a bill, invoice, or receipt, and wherein recalculating the field value comprises:

recalculating at least one of a subtotal field, a tax field, or a total field.

17. The system of claim 10 , wherein detecting the line items block comprises:

heuristically determining geometric bounds of the line items block based on other labeled information blocks; or

using a pre-trained table detection machine learning model.

18. The system of claim 10 , wherein labeling the individual line items comprises:

extracting text in the line items block using optical character recognition;

identifying numeric strings from the extracted text;

determining that the identified numeric strings satisfy arithmetic constraints;

using vertical positions of the numeric strings to define the individual line items; and

labeling the numeric strings as line item amounts and corresponding text as line item description.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2023
From: RIMCHALA, THARATHORN
To: INTUIT INC.
Reel/Frame 063332/0639 →
Cited By (3)
US 12,374,139 US 12,394,237 US 12,547,946