IP Library Granted Patent US 9,183,452
Granted Patent B2
US 9,183,452 · App. 14/269,777 · Granted Nov 10, 2015

Text recognition for textually sparse images

Inventors: Alessandro Bissacco (Los Angeles, CA); Hartmut Neven (Malibu, CA)
Assignee: Google Inc.
G06K9/34G06K9/32G06K9/3208G06K9/3233G06K9/342G06K9/6857G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,183,452
App. No.
14/269,777
Granted
Nov 10, 2015
Kind
B2
Abstract

A text recognition server is configured to recognize text in a sparse text image. Specifically, given an image, the server specifies a plurality of “patches” (blocks of pixels within the image). The system applies a text detection algorithm to the patches to determine a number of the patches that contain text. This application of the text detection algorithm is used both to estimate the orientation of the image and to determine whether the image is textually sparse or textually dense. If the image is determined to be textually sparse, textual patches are identified and grouped into text regions, each of which is then separately processed by an OCR algorithm, and the recognized text for each region is combined into a result for the image as a whole.

Claims (67)

1. A computer-implemented method of recognizing text in an image, comprising:

identifying a first plurality of patches within the image;

applying, to each patch of the first plurality of patches, a first text detection algorithm that indicates whether a patch contains text, thereby identifying a first set of patches that contain text;

determining whether the image represents sparse text or dense text based at least in part on the identified first set of patches;

responsive to the image representing sparse text:

identifying within the image, using a second text detection algorithm, a second set of patches that contain text, and

performing optical character recognition on the second set of patches that contain text to obtain a textual result;

responsive to the image representing dense text:

performing optical character recognition on the image as a whole to produce the textual result; and

storing the textual result.

2. The computer-implemented method of claim 1 , wherein determining that the image represents sparse text comprises:

determining a number of patches within the identified first set of patches that contain text; and

comparing the number of patches that contain text to a threshold, wherein the image is deemed to represent sparse text when the number of patches is below the threshold.

3. The computer-implemented method of claim 1 , wherein determining that the image represents sparse text comprises:

determining an approximate number of characters in the image;

comparing the approximate number of characters to a threshold, wherein the properly-oriented image is deemed to represent sparse text when the number of characters is below the threshold.

4. The computer-implemented method of claim 1 , further comprising:

grouping patches of the second set of patches that contain text and are proximate to one another into a text region,

wherein performing optical character recognition on the second set of patches that contain text to obtain the textual result comprises performing optical character recognition on the text region.

5. The computer-implemented method of claim 4 , wherein grouping patches of the second set of patches that are proximate to one another into the text region comprises combining overlapping patches and forming the text region from a union of areas of the combined patches.

6. The computer-implemented method of claim 1 , wherein the second set of patches comprises patches that overlap with each other or patches of different sizes.

7. A computer system for recognizing text in an image, comprising:

at least one computer processor; and

a computer program executable by the computer processor and performing actions comprising:

identifying a first plurality of patches within the image,

applying, to each patch of the first plurality of patches, a first text detection algorithm that indicates whether a patch contains text, thereby identifying a first set of patches that contain text,

determining whether the image represents sparse text or dense text based at least in part on the identified first set of patches,

responsive to the image representing sparse text:

identifying within the image, using a second text detection algorithm, a second set of patches that contain text,

performing optical character recognition on the second set of patches that contain text to obtain a textual result,

responsive to the image representing dense text:

performing optical character recognition on the image as a whole to produce the textual result, and

storing the textual result.

8. The computer system of claim 7 , wherein determining that the image represents sparse text comprises:

determining a number of patches within the identified first set of patches that contain text; and

comparing the number of patches that contain text to a threshold, wherein the image is deemed to represent sparse text when the number of patches is below the threshold.

9. The computer system of claim 7 , wherein determining that the image represents sparse text comprises:

determining an approximate number of characters in the image;

comparing the approximate number of characters to a threshold, wherein the properly-oriented image is deemed to represent sparse text when the number of characters is below the threshold.

10. The computer system of claim 7 , wherein the actions further comprise:

grouping patches of the second set of patches that contain text and are proximate to one another into a text region,

wherein performing optical character recognition on the second set of patches that contain text to obtain the textual result comprises performing optical character recognition on the text region.

11. The computer system of claim 10 , wherein grouping patches of the second set of patches that are proximate to one another into the text region comprises combining overlapping patches and forming the text region from a union of areas of the combined patches.

12. The computer system of claim 7 , wherein the second set of patches comprises patches that overlap with each other or patches of different sizes.

13. A computer-implemented method of recognizing text in an image, comprising:

identifying a first plurality of patches within the image;

applying, to each patch of the first plurality of patches, a first text detection algorithm that indicates whether a patch contains text, thereby identifying a first set of patches that contain text;

determining whether the image represents sparse text or dense text based at least in part on a number of patches within the identified first set of patches;

when the image represents sparse text:

identifying within the image, using a second text detection algorithm, a second set of patches that contain text, and

performing optical character recognition on the second set of patches that contain text to obtain a textual result, and

when the image represents dense text:

performing optical character recognition on the image as a whole to obtain the textual result; and

storing the textual result.

14. The computer-implemented method of claim 13 , wherein determining whether the image represents sparse text or dense text comprises:

determining a number of patches within the identified first set of patches that contain text; and

comparing the number of patches that contain text to a threshold, wherein the image is deemed to represent sparse text when the number of patches is below the threshold and is deemed to represent dense text when the number of patches is above the threshold.

15. The computer-implemented method of claim 13 , wherein determining whether the image represents sparse text or dense text comprises:

determining an approximate number of characters in the image; and

comparing the approximate number of characters to a threshold, wherein the image is deemed to represent sparse text when the approximate number of characters is below the threshold and is deemed to represent dense text when the approximate number of characters is above the threshold.

16. The computer-implemented method of claim 13 , wherein performing optical character recognition on the second set of patches that contain text to obtain a textual result comprises:

grouping patches of the second set of patches that contain text and are proximate to one another into a text region; and

performing optical character recognition on the text region.

17. The computer-implemented method of claim 16 , wherein grouping patches of the second set of patches that are proximate to one another into the text region comprises combining overlapping patches and forming the text region from a union of areas of the combined patches.

18. The computer-implemented method of claim 13 , wherein the second set of patches comprises patches that overlap with each other or patches of different sizes.

19. The computer-implemented method of claim 1 , wherein identifying the second set of patches comprises identifying the first set of patches.

20. The computer system of claim 7 , wherein identifying the second set of patches comprises identifying the first set of patches.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044334/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2014
From: BISSACCO, ALESSANDRO; NEVEN, HARTMUT
To: GOOGLE INC.
Reel/Frame 032823/0001 →
Continuity (2)
Continuation 12608877 · Oct 29, 2009
Related Publication 20150161465A1 · Jun 11, 2015