IP Library › Granted Patent US 10,970,530
Granted Patent B1
US 10,970,530 · App. 16/189,633 · Granted Apr 6, 2021

Grammar-based automated generation of annotated synthetic form training data for machine learning

Inventors: Amit Adam (Haifa, IL); Oron Anschel (Haifa, IL); Or Perel (Tel Aviv, IL); Gal Sabina Star (Haifa, IL); Omri Ben-Eliezer (Ramat Gan, IL); Hadar Averbuch Elor (Ra'anana, IL); Shai Mazor (Binyamina, IL); Wendy Tse (Seattle, WA); Andrea Olgiati (Gilroy, CA); Rahul Bhotika (Bellevue, WA); Stefano Soatto (Pasadena, CA)
Assignee: Amazon Technologies, Inc.
G06K9/00449G06F40/137G06F40/169G06F40/174G06K9/00469G06K9/6257G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,970,530
App. No.
16/189,633
Granted
Apr 6, 2021
Kind
B1
Abstract

Techniques for grammar-based automated generation of annotated synthetic form training data for machine learning are described. A training data generation engine utilizes a defined grammar to construct a layout for a form, select key-value units to place within the layout, and select attribute variants for the key-value units. The form is rendered and stored at a storage location, where it can be provided along with other similarly-generated forms to be used as training data for a machine learning model.

Claims (91)

1. A computer-implemented method comprising:

generating a plurality of images including representations of a corresponding plurality of forms, wherein for each image of the plurality of images the generating comprises:

determining, based on a defined grammar, a plurality of document sections for the form,

selecting, for each of the plurality of document sections based on the defined grammar, one or more key-value units to be placed in the document section,

selecting, for each of the selected one or more key-value units, one or more style attributes for the key-value unit based on a random value, and

placing the one or more key-value units within the document section with the selected style attributes;

storing the plurality of images to a storage location along with a corresponding plurality of annotations; and

providing the plurality of images and the plurality of annotations to be used to train a machine learning (ML) model.

2. The computer-implemented method of claim 1 , wherein for each image of the plurality of images the generating further comprises:

generating, for at least one of the one or more key-value units of the image, a value to be placed within the at least one key-value unit;

determining, based on a second random value, a characteristic comprising at least one of a location, a font size, a font, or a font style for the value; and

placing the value within the at least one key-value unit according to the determined characteristic.

3. The computer-implemented method of claim 1 , wherein the defined grammar specifies that at least one pair of key-value units that are to be placed adjacent to one another.

4. A computer-implemented method comprising:

generating a plurality of documents, wherein for each document of the plurality of documents the generating comprises:

determining, based on a defined grammar, a plurality of sections for the document,

selecting, for each of the plurality of sections based on the defined grammar, one or more key-value units to be placed in the section, and

selecting, for each of the selected one or more key-value units, one or more style attributes and placing the one or more key-value units within the section; and

storing the plurality of documents to a storage location.

5. The computer-implemented method of claim 4 , wherein the plurality of documents includes representations of forms that are stored as image files.

6. The computer-implemented method of claim 5 , further comprising:

generating a plurality of annotation data structures corresponding to the plurality of documents, each annotation data structure indicating at least locations of the one or more key-value units within the corresponding document; and

storing the plurality of annotation data structures along with the plurality of documents at the storage location.

7. The computer-implemented method of claim 6 , further comprising:

obtaining the plurality of annotation data structures and the plurality of documents from the storage location; and

utilizing the plurality of annotation data structures and the plurality of documents to train a machine learning (ML) model.

8. The computer-implemented method of claim 4 , further comprising:

selecting, based on a random value, the defined grammar from a plurality of defined grammars.

9. The computer-implemented method of claim 4 , further comprising:

for at least one of the one or more key-value units of a document,

obtaining a key;

determining, based on a randomization, one or more style attributes for the key;

obtaining a value;

determining, based on a randomization, one or more style attributes for the value; and

placing the key and the value within the at least one key-value unit according to the one or more style attributes for the key and according to the one or more style attributes for the value.

10. The computer-implemented method of claim 9 , wherein the one or more style attributes for the key include one or more of:

a stride amount between characters of the key;

a font size; or

a font.

11. The computer-implemented method of claim 9 , wherein:

obtaining the key comprises selecting the key from a dictionary of keys; and

obtaining the value comprises generating the value based on the key.

12. The computer-implemented method of claim 4 , wherein:

the selecting of each of the one or more style attributes for at least one of the key-value units is based on a random value;

the one or more style attributes for at least one of the key-value units include at least one of:

a width of a line;

a style of a line;

a color of a line;

a background fill color;

a margin or padding amount;

a position of the key; or

a position of the value.

13. The computer-implemented method of claim 4 , wherein the generating further comprises:

for one of the plurality of documents, determining the plurality of sections for the document includes identifying a hierarchy of sections, wherein a first section of the hierarchy includes a second section of the hierarchy that is to be placed within the first section.

14. A system comprising:

a storage service implemented by a first one or more electronic devices; and

a training data generation engine implemented by a second one or more electronic devices, the training data generation engine including instructions that upon execution cause the training data generation engine to:

generate a plurality of documents, wherein for each document of the plurality of documents, the training data generation engine is to:

determine, based on a defined grammar, a plurality of sections for the document,

select, for each of the plurality of sections, one or more key-value units, and

select, for each of the selected one or more key-value units, one or more style attributes and place the one or more key-value units within the section; and

store the plurality of documents to a storage location provided by the storage service.

15. The system of claim 14 , wherein the plurality of documents includes representations of forms and are stored as image files.

16. The system of claim 15 , wherein the training data generation engine is further to:

generate a plurality of annotation data structures corresponding to the plurality of documents, each annotation data structure indicating at least locations of the one or more key-value units within the document; and

store the plurality of annotation data structures along with the plurality of documents at the storage location.

17. The system of claim 16 , further comprising another service of a same provider network as the training data generation engine and the storage service, the another service including instructions that upon execution cause the another service to:

obtain the plurality of annotation data structures and the plurality of documents from the storage location; and

utilize the plurality of annotation data structures and the plurality of documents to train a machine learning (ML) model.

18. The system of claim 14 , wherein the training data generation engine is further to:

select, according to a second random value, the defined grammar from a plurality of defined grammars.

19. The system of claim 14 , wherein the training data generation engine is further to:

for at least one of the one or more key-value units of a document,

obtain a key;

determine one or more style attributes for the key based on a random value;

obtain a value;

determine one or more style attributes for the value based on a random value; and

place the key and the value within the at least one key-value unit according to the one or more style attributes for the key and according to the one or more style attributes for the value.

20. The system of claim 19 , wherein:

for the at least one key-value unit, the one or more style attributes for the key include one or more of:

a stride amount between characters of the key;

a font size; or

a font; and

the one or more style attributes for at the least one key-value unit include at least one of:

a width of a line;

a style of a line;

a color of a line;

a background fill color;

a margin or padding amount;

a position of the key; or

a position of the value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2018
From: ADAM, AMIT; ANSCHEL, ORON; PEREL, OR; STAR, GAL SABINA; BEN-ELIEZER, OMRI; ELOR, HADAR AVERBUCH; MAZOR, SHAI; TSE, WENDY; OLGIATI, ANDREA; BHOTIKA, RAHUL; SOATTO, STEFANO
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 047504/0092 →
Cited By (3)
US 12,260,301 US 12,394,230 US 12,449,789