IP Library › Granted Patent US 10,354,161
Granted Patent B2
US 10,354,161 · App. 15/614,298 · Granted Jul 16, 2019

Detecting font size in a digital image

Inventors: Peijun Chiang (Mountain View, CA); Vijay Yellapragada (Mountain View, CA)
Assignee: INTUIT, INC.
G06K9/325G06K9/00442G06K9/2081G06K9/344G06K9/46G06K9/4604G06K9/48G06T7/11G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,354,161
App. No.
15/614,298
Granted
Jul 16, 2019
Kind
B2
Abstract

The present disclosure relates to optical character recognition, and more specifically techniques for detecting font size in a digital image. Accordingly to one embodiment, a client device receives a digital image of a document having one or more textual components. The client device finds one or more contours bounding the one or more textual components in the digital image of the document. The client device detects a font size for text contained in the digital image using the one or more contours. The client device extracts the text from the digital image upon detecting that the detected font size is above a defined threshold value.

Claims (58)

1. A computer-implemented method for performing optical character recognition on a digital image of a document, the method comprising:

receiving a digital image of a document having one or more textual components;

finding one or more contours bounding the one or more textual components in the digital image of the document;

detecting a font size for text contained in the digital image using the one or more contours, wherein detecting the font size comprises:

creating one or more bounding rectangles encompassing the one or more contours;

calculating an area bound by each contour of the one or more contours; and

for each respective textual component of the one or more textual components bound by a respective contour of the one or more contours, computing a height for the respective textual component by dividing a corresponding area for the respective contour by a width of a corresponding bounding rectangle of the one or more bounding rectangles; and

extracting the text from the digital image upon detecting that the detected font size is above a defined threshold value.

2. The method of claim 1 , further comprising:

adjusting a height of each bounding rectangle of the one or more bounding rectangles resulting in one or more adjusted bounding rectangles based on a height of each respective textual component of the one or more textual components contained within a corresponding contour of the one or more contours.

3. The method of claim 2 , further comprising:

applying a filter to the one or more adjusted bounding rectangles to exclude adjusted bounding rectangles whose corresponding heights are outside a certain range resulting in a list of remaining adjusted bounding rectangles;

calculating a mean height of the remaining adjusted bounding rectangles; and

comparing the mean height with one or more subsequent mean heights calculated for one or more smaller chunks of the digital image to detect the font size for the text contained in the digital image.

4. The method of claim 1 , further comprising:

prompting a user to perform one or more actions to capture one or more subsequent images of the document upon determining that the detected font size is below the defined threshold value.

5. The method of claim 4 , wherein the one or more actions include adjusting a perspective of a camera of a computing device.

6. The method of claim 4 , wherein the one or more actions include capturing a plurality of images, each image representing a different segment of the document.

7. The method of claim 1 , wherein the digital image comprises a video frame.

8. A system comprising:

a processor;

a memory comprising instructions which, when executed on the processor, performs an operation for performing optical character recognition on a digital image of a document, the operation comprising:

receiving the digital image of the document having one or more textual components;

finding one or more contours bounding the one or more textual components in the digital image of the document;

detecting a font size for text contained in the digital image using the one or more contours, wherein detecting the font size comprises:

creating one or more bounding rectangles encompassing the one or more contours;

calculating an area bound by each contour of the one or more contours; and

for each respective textual component of the one or more textual components bound by a respective contour of the one or more contours, computing a height for the respective textual component by dividing a corresponding area for the respective contour by a width of a corresponding bounding rectangle of the one or more bounding rectangles; and

extracting the text from the digital image upon detecting that the detected font size is above a defined threshold value.

9. The system of claim 8 , further comprising:

adjusting a height of each bounding rectangle of the one or more bounding rectangles resulting in one or more adjusted bounding rectangles based on a height of each respective textual component of the one or more textual components contained within a corresponding contour of the one or more contours.

10. The system of claim 9 , further comprising:

applying a filter to the one or more adjusted bounding rectangles to exclude adjusted bounding rectangles whose corresponding heights are outside a certain range resulting in a list of remaining adjusted bounding rectangles;

calculating a mean height of the remaining adjusted bounding rectangles; and

comparing the mean height with one or more subsequent mean heights calculated for one or more smaller chunks of the digital image to detect the font size for the text contained in the digital image.

11. The system of claim 8 , further comprising:

prompting a user to perform one or more actions to capture one or more subsequent images of the document upon determining that the detected font size is below the defined threshold value.

12. The system of claim 11 , wherein the one or more actions include adjusting a perspective of a camera of a computing device.

13. The system of claim 11 , wherein the one or more actions include capturing a plurality of images, each image representing a different segment of the document.

14. A non-transitory computer-readable medium comprising instructions which, when executed on one or more processors, performs an operation for performing optical character recognition on a digital image of a document, the operation comprising:

receiving the digital image of the document having one or more textual components;

finding one or more contours bounding the one or more textual components in the digital image of the document;

detecting a font size for text contained in the digital image using the one or more contours, wherein detecting the font size comprises:

creating one or more bounding rectangles encompassing the one or more contours;

calculating an area bound by each contour of the one or more contours; and

for each respective textual component of the one or more textual components bound by a respective contour of the one or more contours, computing a height for the respective textual component by dividing a corresponding area for the respective contour by a width of a corresponding bounding rectangle of the one or more bounding rectangles; and

extracting the text from the digital image upon detecting that the detected font size is above a defined threshold value.

15. The non-transitory computer-readable medium of claim 14 , further comprising:

adjusting a height of each bounding rectangle of the one or more bounding rectangles resulting in one or more adjusted bounding rectangles based on a height of each respective textual component of the one or more textual components contained within a corresponding contour of the one or more contours.

16. The non-transitory computer-readable medium of claim 15 , further comprising:

applying a filter to the one or more adjusted bounding rectangles to exclude adjusted bounding rectangles whose corresponding heights are outside a certain range resulting in a list of remaining adjusted bounding rectangles;

calculating a mean height of the remaining adjusted bounding rectangles; and

comparing the mean height with one or more subsequent mean heights calculated for one or more smaller chunks of the digital image to detect the font size for the text contained in the digital image.

17. The non-transitory computer-readable medium of claim 14 , further comprising:

prompting a user to perform one or more actions to capture one or more subsequent images of the document upon determining that the detected font size is below the defined threshold value.

18. The non-transitory computer-readable medium of claim 17 , wherein the one or more actions include adjusting a perspective of a camera of a computing device.

19. The non-transitory computer-readable medium of claim 17 , wherein the one or more actions include capturing a plurality of images, each image representing a different segment of the document.

20. The non-transitory computer-readable medium of claim 14 , wherein the digital image comprises a video frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2017
From: CHIANG, PEIJUN; YELLAPRAGADA, VIJAY
To: INTUIT, INC.
Reel/Frame 042602/0741 →
Continuity (1)
Related Publication 20180349722A1 · Dec 6, 2018