IP Library › Granted Patent US 11,816,710
Granted Patent B2
US 11,816,710 · App. 17/653,097 · Granted Nov 14, 2023

Identifying key-value pairs in documents

Inventors: Yang Xu (San Jose, CA); Jiang Wang (Santa Clara, CA); Shengyang Dai (Dublin, CA)
Assignee: Google LLC
G06Q30/04G06V30/412G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,816,710
App. No.
17/653,097
Granted
Nov 14, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for converting unstructured documents to structured key-value pairs. In one aspect, a method includes: providing an image of a document to a detection model, wherein: the detection model is configured to process the image to generate an output that defines one or more bounding boxes generated for the image; and each bounding box generated for the image is predicted to enclose a key-value pair including key textual data and value textual data, wherein the key textual data defines a label that characterizes the value textual data; and for each of the one or more bounding boxes generated for the image: identifying textual data enclosed by the bounding box using an optical character recognition technique; and determining whether the textual data enclosed by the bounding box defines a key-value pair.

Claims (58)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving a request requesting the data processing hardware to characterize key-value pairs within an image, the image comprising a bounding box enclosing textual data within the image;

determining whether a first portion of the textual data enclosed by the bounding box includes a key from a predetermined set of valid keys; and

when the first portion of the textual data enclosed by the bounding box includes the key from the predetermined set of valid keys:

identifying a second portion of the textual data enclosed by the bounding box, the second portion comprising non-key text of the textual data;

identifying one or more valid types of non-key text associated with the key;

determining whether a type of the second portion of the textual data is included in the one or more valid types of non-key text associated with the key;

when the type of the second portion of the textual data is included in the one or more valid types of non-key text associated with the key, determining that the textual data enclosed by the bounding box defines a valid key-value pair; and

providing the valid key-value pair.

2. The method of claim 1 , wherein the operations further comprise providing the image to a detection model.

3. The method of claim 2 , wherein the operations further comprise:

receiving a document from a user; and

converting the document to the image, wherein the image depicts the document.

4. The method of claim 2 , wherein providing the image to the detection model comprises:

identifying a particular class of the image;

training the detection model on the particular class; and

providing the image to the detection model trained to process images of the particular class.

5. The method of claim 2 , wherein the detection model comprises a neural network model.

6. The method of claim 5 , wherein:

the neural network model is trained on a set of training examples, each training example comprising a training input and a target output;

the training input comprises a training image; and

the target output comprises data defining one or more bounding boxes in the training image that each enclose a respective valid key-value pair.

7. The method of claim 2 , wherein determining whether the first portion of the textual data enclosed by the bounding box includes the key from the predetermined set of valid keys further comprises, when the first portion of the textual data enclosed by the bounding box does not include the key from the predetermined set of valid keys, determining that the textual data enclosed by the bounding box does represents an invalid key-value pair.

8. The method of claim 1 , wherein identifying the one or more valid types of the non-key text of the textual data associated with the key comprises mapping the key to the one or more valid types for values corresponding to the key using a predetermined mapping.

9. The method of claim 8 , wherein the predetermined set of valid keys and the predetermined mapping are provided by a user.

10. The method of claim 1 , wherein determining whether the second portion of the textual data is included in the one or more valid types of non-key text associated with the key comprises:

mapping the key to a temporal type of the second portion of the textual data; and

indicating that a value corresponding to the key should have a temporal data type.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a request requesting the data processing hardware to characterize key-value pairs within an image, the image comprising a bounding box enclosing textual data within the image;

determining whether a first portion of the textual data enclosed by the bounding box includes a key from a predetermined set of valid keys; and

when the first portion of the textual data enclosed by the bounding box includes the key from the predetermined set of valid keys:

identifying a second portion of the textual data enclosed by the bounding box, the second portion comprising non-key text of the textual data;

identifying one or more valid types of non-key text associated with the key;

determining whether a type of the second portion of the textual data is included in the one or more valid types of non-key text associated with the key;

when the type of the second portion of the textual data is included in the one or more valid types of non-key text associated with the key, determining that the textual data enclosed by the bounding box defines a valid key-value pair; and

providing the valid key-value pair.

12. The system of claim 11 , wherein the operations further comprise providing the image to a detection model.

13. The system of claim 12 , wherein the operations further comprise:

receiving a document from a user; and

converting the document to the image, wherein the image depicts the document.

14. The system of claim 12 , wherein providing the image to the detection model comprises:

identifying a particular class of the image;

training the detection model on the particular class; and

providing the image to the detection model trained to process images of the particular class.

15. The system of claim 12 , wherein the detection model comprises a neural network model.

16. The system of claim 15 , wherein:

the neural network model is trained on a set of training examples, each training example comprising a training input and a target output;

the training input comprises a training image; and

the target output comprises data defining one or more bounding boxes in the training image that each enclose a respective valid key-value pair.

17. The system of claim 12 , wherein determining whether the first portion of the textual data enclosed by the bounding box includes they key from the predetermined set of valid keys further comprises, when the first portion of the textual data enclosed by the bounding box does not include the key from the predetermined set of valid keys, determining that the textual data enclosed by the bounding box does represents an invalid key-value pair.

18. The system of claim 11 , wherein identifying the one or more valid types of the non-key text of the textual data associated with the key comprises mapping the key to the one or more valid types for values corresponding to the key using a predetermined mapping.

19. The system of claim 18 , wherein the predetermined set of valid keys and the predetermined mapping are provided by a user.

20. The system of claim 11 , wherein determining whether the second portion of the textual data is included in the one or more valid types of non-key text associated with the key comprises:

mapping the key to a temporal type of the second portion of the textual data; and

indicating that a value corresponding to the key should have a temporal data type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2022
From: XU, YANG; WANG, JIANG; DAI, SHENGYAN
To: GOOGLE LLC
Reel/Frame 060330/0808 →
Continuity (3)
Continuation 16802864 · Feb 27, 2020
Provisional Application 62811331 · Feb 27, 2019
Related Publication 20220309549A1 · Sep 29, 2022