IP Library › Granted Patent US 10,210,384
Granted Patent B2
US 10,210,384 · App. 15/218,907 · Granted Feb 19, 2019

Optical character recognition (OCR) accuracy by combining results across video frames

Inventors: Vijay Yellapragada (Mountain View, CA); Peijun Chiang (Mountain View, CA); Sreeneel K. Maddika (Mountain View, CA)
Assignee: INTUIT INC.
G06K9/00442G06F17/2765G06K9/00469G06Q40/123G06T7/0081G06K2209/01G06T2207/10004G06T2207/30176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,210,384
App. No.
15/218,907
Granted
Feb 19, 2019
Kind
B2
Abstract

The present disclosure relates to optical character recognition using captured video. According to one embodiment, using a first image in stream of images depicting a document, the device extracts text data in a portion of the document depicted in the first image and determines a first confidence level regarding an accuracy of the extracted text data. If the first confidence level satisfies a threshold value, the device saves the extracted text data as recognized content of the source document. Otherwise, the device extracts the text data from the portion of the document as depicted in one or more second images in the stream and determines a second confidence level for the text data extracted from each second image until identifying one of the second images where the second confidence level associated with the text data extracted from the identified second image satisfies the threshold value.

Claims (58)

1. A method for evaluating text depicted in images of a source document, comprising:

receiving a stream of digital images depicting the source document, wherein the document comprises a semi-structured document with a plurality of fields;

for a first digital one of the images:

extracting text data in a portion of the document as depicted in the first image, and

determining a first confidence level regarding an accuracy of the extracted text data;

upon determining that the first confidence level satisfies a threshold value, saving the extracted text data as recognized content of the source document;

upon determining that the first confidence level does not satisfy the threshold value:

extracting the text data from the portion of the document as depicted in one or more second images in the stream of digital images, and

determining a second confidence level for the text data extracted from each second image until identifying one of the second images such that the second confidence level associated with the text data extracted from the identified second image satisfies the threshold value;

evaluating text extracted from the source document for a first one of the fields across the first digital image and the one or more second digital images; and

identifying an extracted text as a most likely value of the first field based, at least in part, on a frequency in which the identified extracted text was extracted from the first one of the fields and an average confidence level associated with the identified extracted text.

2. The method of claim 1 , wherein extracting text data from the document comprises associating an extracted name of a field with data entered in the field.

3. An apparatus, comprising:

a processor; and

a memory having instructions which, when executed by the processor, performs an operation for evaluating text depicted in images of a source document, the operation comprising:

receiving a stream of digital images depicting the source document, wherein the document comprises a semi-structured document with a plurality of fields;

for a first digital one of the images:

extracting text data in a portion of the document as depicted in the first image, and

determining a first confidence level regarding an accuracy of the extracted text data;

upon determining that the first confidence level satisfies a threshold value, saving the extracted text data as recognized content of the source document;

upon determining that the first confidence level does not satisfy the threshold value:

extracting the text data from the portion of the document as depicted in one or more second images in the stream of digital images, and

determining a second confidence level for the text data extracted from each second image until identifying one of the second images such that the second confidence level associated with the text data extracted from the identified second image satisfies the threshold value;

evaluating text extracted from the source document for a first one of the fields across the first digital image and the one or more second digital images; and

identifying an extracted text as a most likely value of the first field based, at least in part, on a frequency in which the identified extracted text was extracted from the first one of the fields and an average confidence level associated with the identified extracted text.

4. The apparatus of claim 3 , wherein the operation further comprises:

identifying, in the semi-structured document, one or more key data fields, wherein the extracting text data from the document comprises extracting text data from at least the one or more key data fields.

5. A non-transitory computer-readable medium comprising instructions which, when executed on one or more processors, performs an operation for evaluating text depicted in images of a source document, the operation comprising:

receiving a stream of digital images depicting the source document, wherein the document comprises a semi-structured document with a plurality of fields;

for a first digital one of the images:

extracting text data in a portion of the document as depicted in the first image, and

determining a first confidence level regarding an accuracy of the extracted text data;

upon determining that the first confidence level satisfies a threshold value, saving the extracted text data as recognized content of the source document;

upon determining that the first confidence level does not satisfy the threshold value:

extracting the text data from the portion of the document as depicted in one or more second images in the stream of digital images, and

determining a second confidence level for the text data extracted from each second image until identifying one of the second images such that the second confidence level associated with the text data extracted from the identified second image satisfies the threshold value;

evaluating text extracted from the source document for a first one of the fields across the first digital image and the one or more second digital images; and

identifying an extracted text as a most likely value of the first field based, at least in part, on a frequency in which the identified extracted text was extracted from the first one of the fields and an average confidence level associated with the identified extracted text.

6. The non-transitory computer-readable medium of claim 5 , wherein the operation further comprises:

identifying, in the semi-structured document, one or more key data fields, wherein the extracting text data from the document comprises extracting text data from at least the one or more key data fields.

7. The method of claim 1 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

correcting the extracted text data for the first field according to a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data based on corrections performed on the extracted text data to cause the extracted text to conform to the rule.

8. The method of claim 1 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

determining that the format of the extracted text data does not match a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data to a lower confidence level in response to the determination.

9. The system of claim 3 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

correcting the extracted text data for the first field according to a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data based on corrections performed on the extracted text data to cause the extracted text to conform to the rule.

10. The system of claim 3 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

determining that the format of the extracted text data does not match a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data to a lower confidence level in response to the determination.

11. The non-transitory computer-readable medium of claim 5 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

correcting the extracted text data for the first field according to a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data based on corrections performed on the extracted text data to cause the extracted text to conform to the rule.

12. The non-transitory computer-readable medium of claim 5 , wherein determining the first confidence level regarding an accuracy of the extracted text data comprises:

determining that the format of the extracted text data does not match a rule defining a format of a permitted text for the first field; and

adjusting the confidence level associated with the extracted text data to a lower confidence level in response to the determination.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2016
From: YELLAPRAGADA, VIJAY; CHIANG, PEIJUN; MADDIKA, SREENEEL K.
To: INTUIT INC.
Reel/Frame 039248/0377 →
Continuity (1)
Related Publication 20180025222A1 · Jan 25, 2018
Cited By (2)
US 12,306,850 US 12,482,285