IP Library Granted Patent US 10,867,171
Granted Patent B1
US 10,867,171 · App. 16/167,334 · Granted Dec 15, 2020

Systems and methods for machine learning based content extraction from document images

Inventors: Alexander Wesley Contryman (San Francisco, CA); Jacob Ryan van Gogh (Mountain View, CA); Manu Shukla (Sunnyvale, CA)
Assignee: OMNISCIENCE CORPORATION
G06K9/00463G06K9/00456G06N20/20G06T3/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,867,171
App. No.
16/167,334
Granted
Dec 15, 2020
Kind
B1
Abstract

A method and apparatus for recognizing and extracting data from a form depicted within an image of a document are described. The method may include receiving the image of the document, the image depicting the form and data contained one the form. The method may also include transforming the image of the document to a set of one or more key, value pairs by processing the image of the document with a sequence of two or more trained machine learning based image analysis processes, wherein keys are relevant to forms of the type depicted in the form, and wherein each value is associated with a key. The method may also include generating a data output that comprises the set of key, value pairs for textual data recognized and extracted from the form depicted in the image.

Claims (47)

1. A method for recognizing and extracting data from a form depicted within an image of a document, the method comprising:

receiving, by a computer processing system, the image of the document, the image depicting the form and data contained on the form;

transforming, by the computer processing system, the image of the document to a set of one or more key, value pairs by processing the image of the document with a sequence of two or more trained machine learning based image analysis processes, wherein keys are relevant to forms of the type depicted in the form, and wherein each value is associated with a key, wherein the two or more trained machine learning based image analysis processes comprise a plurality of different machine learning based image analysis processes performed within a document processing pipeline; and

generating, by the computer processing system, a data output that comprises the set of key, value pairs for textual data recognized and extracted from the form depicted in the image.

2. The method of claim 1 , wherein a description of a layout and a structure of the form depicted in the image of the document is unknown to the computer processing system.

3. The method of claim 1 , wherein the method further comprises:

performing a first machine learning based image analysis process to predict a rotation of the form depicted in the image of the document based on features extracted by a first machine learning based image analysis of the received image of the document; and

generating a rotated document image that rotates the form depicted in the image of the document based on the predicted rotation.

4. The method of claim 3 , wherein the form is rotated to a vertical orientation as depicted in the rotated document image.

5. The method of claim 3 , wherein the method further comprises:

performing a second machine learning based image analysis process to detect and locate text segments on the form depicted in the rotated document image based on features extracted by a second machine learning based image analysis of the rotated document image.

6. The method of claim 5 , wherein the method further comprises:

performing a third machine learning based image analysis process to extract features in the detected and located text segments based using a third machine learning based image analysis of the detected and located text segments;

performing a fourth machine learning based image analysis process to generate text character probabilities for each text character within the detected and located text segments based using a fourth machine learning based image analysis of the features extracted from the detected and located text segments by the third machine learning based image analysis;

selecting a most likely text character for each text character within the detected and located text segments based on the generated text character probabilities; and

generating textual content for each detected and located text segment based on the selection of text characters for said detected and located text segment.

7. The method of claim 6 , further comprising:

performing a fifth machine learning based image analysis process to classify textual content for detected and located text segments using a fifth machine learning based image analysis of at least a location associated with a segment and text content contained within said segment.

8. The method of claim 1 , wherein each of the plurality of different machine learning based image analysis processes comprises a machine learning analysis system executed by the computer processing system, and wherein each machine learning analysis system corresponding to each of the plurality of different machine learning based image analysis processes is trained using different training data relevant to a phase in the document processing pipeline in which machine learning image analysis is being performed.

9. The method of claim 1 , wherein the form depicted within the image of the document comprises one of a medical form, an insurance form, or a loan application form.

10. A non-transitory computer readable storage medium including instructions that, when executed by a computer processing system, cause the computer processing system to perform operations for recognizing and extracting data from a form depicted within an image of a document, the operations comprising:

receiving the image of the document, the image depicting the form and data contained one the form, and wherein the form depicted in the image is unstructured;

transforming the image of the document to a set of one or more key, value pairs by processing the image of the document with a sequence of two or more trained machine learning based image analysis processes, wherein keys are relevant to forms of the type depicted in the image, and wherein each value is associated with a key, wherein the two or more trained machine learning based image analysis processes comprise a plurality of different machine learning based image analysis processes performed within a document processing pipeline; and

generating a data output that comprises the set of key, value pairs for textual data recognized and extracted from the form depicted in the image.

11. The non-transitory computer readable storage medium of claim 10 , wherein a description of a layout and a structure of the form depicted in the image of the document is unknown to the computer processing system.

12. The non-transitory computer readable storage medium of claim 10 , wherein the operations further comprise:

performing a first machine learning based image analysis process to predict a rotation of the form depicted in the image of the document based on features extracted by a first machine learning based image analysis of the received image of the document; and

generating a rotated document image that rotates the form depicted in the image of the document based on the predicted rotation.

13. The non-transitory computer readable storage medium of claim 12 , wherein the form is rotated to a vertical orientation as depicted in the rotated document image.

14. The non-transitory computer readable storage medium of claim 12 , wherein the operations further comprise:

performing a second machine learning based image analysis process to detect and locate text segments on the form depicted in the rotated document image based on features extracted by a second machine learning based image analysis of the rotated document image.

15. The non-transitory computer readable storage medium of claim 14 , wherein the method further comprises:

performing a third machine learning based image analysis process to extract features in the detected and located text segments based using a third machine learning based image analysis of the detected and located text segments;

performing a fourth machine learning based image analysis process to generate text character probabilities for each text character within the detected and located text segments based using a fourth machine learning based image analysis of the features extracted from the detected and located text segments by the third machine learning based image analysis;

selecting a most likely text character for each text character within the detected and located text segments based on the generated text character probabilities; and

generating textual content for each detected and located text segment based on the selection of text characters for said each detected and located text segment.

16. The non-transitory computer readable storage medium of claim 15 , wherein the operations further comprise:

performing a fifth machine learning based image analysis process to classify textual content for detected and located text segments using a fifth machine learning based image analysis of at least a location associated with a segment and text content contained within said segment.

17. The non-transitory computer readable storage medium of claim 10 , wherein each of the plurality of different machine learning based image analysis processes comprises of techniques for machine learning analysis system executed by the computer processing system, and wherein each machine learning analysis system corresponding to each of the plurality of different machine learning based image analysis processes is trained using different training data relevant to a phase in the document processing pipeline.

18. The non-transitory computer readable storage medium of claim 10 , wherein the form depicted within the image of the document comprises one of a medical form, an insurance form, or a loan application form.

19. A system, comprising:

a network interface that receives an image of a document, the image depicting a form and data contained on the form, and wherein the form depicted in the image is unstructured;

a memory that stores the image of the document; and

a processor coupled with the memory configured to access the image of the document and further configured to:

transform the image of the document to a set of one or more key, value pairs by processing the image of the document with a sequence of two or more trained machine learning based image analysis processes, wherein keys are relevant to forms of the type depicted in the image, and wherein each value is associated with a key, wherein the two or more trained machine learning based image analysis processes comprise a plurality of different machine learning based image analysis processes performed within a document processing pipeline, and

generate a data output that comprises the set of key, value pairs for textual data recognized and extracted from the form depicted in the image.

20. The system of claim 19 , wherein a description of a layout and a structure of the form depicted in the image of the document is unknown to the computer processing system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2024
From: OMNISCIENCE (ABC), LLC
To: OMNISCIENCE STRATEGIES CORPORATION
Reel/Frame 067937/0867 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2019
From: CONTRYMAN, ALEXANDER WESLEY; VAN GOGH, JACOB RYAN; SHUKLA, MANU
To: OMNISCIENCE CORP.
Reel/Frame 048941/0843 →
Cited By (13)
US 12,307,369 US 12,340,319 US 12,354,022 US 12,423,949 US 12,482,286 US 12,525,048 US 12,541,989 US 12,585,977 US 12,586,402 US 12,608,743 US 12,639,362 US 12,688,667 US 12,718,605