IP Library › Granted Patent US 10,740,638
Granted Patent B1
US 10,740,638 · App. 15/858,297 · Granted Aug 11, 2020

Data element profiles and overrides for dynamic optical character recognition based data extraction

Inventor: David Annis (Edmond, OK)
Assignee: BUSINESS IMAGING SYSTEMS, INC.
G06K9/2063G06K9/626G06K9/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,740,638
App. No.
15/858,297
Granted
Aug 11, 2020
Kind
B1
Abstract

A method for dynamic optical character recognition based data extraction includes: analyzing an image; detecting a first identifier associated with a first content type in an image; providing a first data extraction model for the first content type, the first data extraction model including definitions for a plurality of data types; performing an optical character recognition pass on the image to identify a plurality of characters of the image; and extracting a set of data elements from the image based on the first data extraction model and the plurality of characters of the image identified by performing the optical character recognition pass on the image.

Claims (74)

1. A method for dynamic optical character recognition based data extraction, comprising:

analyzing an image;

detecting a first identifier associated with a first content type in an image;

providing a first data extraction model for the first content type, the first data extraction model including definitions for a plurality of data types;

performing an optical character recognition pass on the image to identify a plurality of characters of the image; and

extracting a set of data elements from the image based on the first data extraction model and the plurality of characters of the image identified by performing the optical character recognition pass on the image, the set of data elements corresponding to a chosen expression pattern, the chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the set of data elements including a plurality of distinct items to be extracted by optical character recognition from the image, the plurality of distinct items corresponding to the chosen expression pattern.

2. The method of claim 1 , further comprising:

analyzing a second image;

detecting a second identifier associated with a second content type in the second image;

providing a second data extraction model for the second content type, the second data extraction model including definitions for a second plurality of data types;

performing an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extracting a second set of data elements from the second image based on the second data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image, the second set of data elements corresponding to a second chosen expression pattern, the second chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the second set of data elements including a second plurality of distinct items to be extracted from the image corresponding to the second chosen expression pattern.

3. The method of claim 2 , wherein the first data extraction model and the second data extraction model are stored in a memory.

4. The method of claim 1 , further comprising:

analyzing a second image;

detecting a second identifier associated with a second content type in the second image;

applying a rule change to modify the first data extraction model based on the second identifier associated with the second content type;

performing an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extracting a second set of data elements from the second image based on a modified version of the first data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image.

5. The method of claim 4 , wherein the rule change is based on a pre-programmed override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

6. The method of claim 4 , wherein the rule change is based upon a user-specified override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

7. The method of claim 6 , further comprising:

presenting an alert when the second identifier associated with the second content type is detected, the alert prompting a user to determine whether the first data extraction model is valid for the second content type; and

receiving a user input including the user-specified override for the first data extraction model.

8. A system for dynamic optical character recognition based data extraction, comprising:

a controller including at least one processor configured to execute one or more modules stored by a memory that is communicatively coupled to the at least one processor, the one or more modules, when executed, causing the processor to:

analyze an image;

detect a first identifier associated with a first content type in an image;

provide a first data extraction model for the first content type, the first data extraction model including definitions for a plurality of data types;

perform an optical character recognition pass on the image to identify a plurality of characters of the image; and

extract a set of data elements from the image based on the first data extraction model and the plurality of characters of the image identified by performing the optical character recognition pass on the image, the set of data elements corresponding to a chosen expression pattern, the chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the set of data elements including a plurality of distinct items to be extracted by optical character recognition from the image, the plurality of distinct items corresponding to the chosen expression pattern.

9. The system of claim 8 , wherein the one or more modules, when executed, cause the processor to:

analyze a second image;

detect a second identifier associated with a second content type in the second image;

provide a second data extraction model for the second content type, the second data extraction model including definitions for a second plurality of data types;

perform an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extract a second set of data elements from the second image based on the second data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image, the second set of data elements corresponding to a second chosen expression pattern, the second chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the second set of data elements including a second plurality of distinct items corresponding to the second chosen expression pattern.

10. The system of claim 9 , wherein the first data extraction model and the second data extraction model are stored in the memory.

11. The system of claim 8 , wherein the one or more modules, when executed, cause the processor to:

analyze a second image;

detect a second identifier associated with a second content type in the second image;

apply a rule change to modify the first data extraction model based on the second identifier associated with the second content type;

perform an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extract a second set of data elements from the second image based on a modified version of the first data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image, the second set of data elements corresponding to a second chosen expression pattern, the second chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the second set of data elements including a second plurality of distinct items corresponding to the second chosen expression pattern.

12. The system of claim 11 , wherein the rule change is based on a pre-programmed override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

13. The system of claim 11 , wherein the rule change is based upon a user-specified override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

14. The system of claim 13 , wherein the one or more modules, when executed, cause the processor to:

present an alert, via a display communicatively coupled to the processor, when the second identifier associated with the second content type is detected, the alert prompting a user to determine whether the first data extraction model is valid for the second content type; and

receive a user input, via an input device communicatively coupled to the processor, the user input including the user-specified override for the first data extraction model.

15. A system for dynamic optical character recognition based data extraction, comprising:

an imaging device;

a controller in communication with the imaging device, the controller including at least one processor configured to execute one or more modules stored by a memory that is communicatively coupled to the at least one processor, the one or more modules, when executed, causing the processor to:

analyze an image received from the imaging device;

detect a first identifier associated with a first content type in an image;

provide a first data extraction model for the first content type, the first data extraction model including definitions for a plurality of data types;

perform an optical character recognition pass on the image to identify a plurality of characters of the image; and

extract a set of data elements from the image based on the first data extraction model and the plurality of characters of the image identified by performing the optical character recognition pass on the image, the set of data elements corresponding to a chosen expression pattern, the chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the set of data elements including a plurality of distinct items to be extracted by optical character recognition from the image, the plurality of distinct items corresponding to the chosen expression pattern.

16. The system of claim 15 , wherein the one or more modules, when executed, cause the processor to:

analyze a second image received from the imaging device;

detect a second identifier associated with a second content type in the second image;

provide a second data extraction model for the second content type, the second data extraction model including definitions for a second plurality of data types;

perform an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extract a second set of data elements from the second image based on the second data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image, the second set of data elements corresponding to a second chosen expression pattern, the second chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the second set of data elements including a second plurality of distinct items corresponding to the second chosen expression pattern.

17. The system of claim 15 , wherein the one or more modules, when executed, cause the processor to:

analyze a second image received from the imaging device;

detect a second identifier associated with a second content type in the second image;

apply a rule change to modify the first data extraction model based on the second identifier associated with the second content type;

perform an optical character recognition pass on the second image to identify a plurality of characters of the second image; and

extract a second set of data elements from the second image based on a modified version of the first data extraction model and the plurality of characters of the second image identified by performing the optical character recognition pass on the second image, the second set of data elements corresponding to a second chosen expression pattern, the second chosen expression pattern including a plurality of at least one of numbers and letters arranged in a chosen format, the second set of data elements including a second plurality of distinct items corresponding to the second chosen expression pattern.

18. The system of claim 17 , wherein the rule change is based on a pre-programmed override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

19. The system of claim 17 , wherein the rule change is based upon a user-specified override including a data element profile for the first data extraction model, wherein the data element profile is associated with at least one data element in relation to the second content type.

20. The system of claim 19 , wherein the one or more modules, when executed, cause the processor to:

present an alert, via a display communicatively coupled to the processor, when the second identifier associated with the second content type is detected, the alert prompting a user to determine whether the first data extraction model is valid for the second content type; and

receive a user input, via an input device communicatively coupled to the processor, the user input including the user-specified override for the first data extraction model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2018
From: ANNIS, DAVID
To: BUSINESS IMAGING SYSTEMS, INC.
Reel/Frame 046697/0208 →
Continuity (1)
Provisional Application 62440777 · Dec 30, 2016
Cited By (6)
US 12,243,337 US 12,259,930 US 12,260,662 US 12,268,478 US 12,288,406 US 12,564,339