IP Library › Granted Patent US 9,852,348
Granted Patent B2
US 9,852,348 · App. 14/690,274 · Granted Dec 26, 2017

Document scanner

Inventors: Krishnendu Chaudhury (Saratoga, CA); Lu Chen (Chapel Hill, NC); David Petrou (Brooklyn, NY); Blaise Aguera-Arcas (Seattle, WA)
Assignee: Google LLC
G06K9/18G06K9/00463G06K9/00483G06K9/3275G06K9/342G06K9/4642G06K9/6211G06T3/4038G06T11/60G06K2009/363
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,852,348
App. No.
14/690,274
Granted
Dec 26, 2017
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, to generate a scannable document. In one aspect, a method includes receiving a scan request, wherein the scan request includes a plurality of text images; for each text image of the plurality of text images: rectifying the text image to generate a text image with parallel image lines, generating a plurality of word bounding boxes that enclose one or more connected components in the text image, wherein each word bounding box is associated with a respective word, and generating, for each respective word in the text image, a plurality of points that represent the respective word; combining the plurality of text images to form a single text document; and providing the combined image as a scannable document.

Claims (72)

1. A computer implemented method, the method comprising:

receiving a scan request, wherein the scan request includes a plurality of text images, each text image representing a portion of a text document, wherein the plurality of text images includes a first text image and a second text image that at least partially overlap;

for each text image of the plurality of text images:

rectifying the text image to generate a text image with parallel image lines,

generating a plurality of word bounding boxes that enclose one or more connected components in the text image, wherein each word bounding box is associated with a respective word of the text image, and

generating, for each respective word bounding box in the text image, a vector that describes a shape of the respective word;

combining text images of the plurality of text images that at least partially overlap to form a single text document including combining the first text image of the plurality of text images and the second text image of the plurality of text images by matching one or more shape descriptors from a first set of vectors generated for each respective word in the first text image and one or more shape descriptors from a second set of vectors generated for each respective word in the second text image; and

providing the combined image as a scannable document.

2. The method of claim 1 , wherein rectifying each text image of the plurality of text images comprises:

determining a plurality of connected components in the text image, each connected component being a filled portion of a symbol;

generating a plurality of image lines in the text image including vertical linelets and horizontal linelets based on the plurality of connected components;

calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines; and

applying a geometric formula to the first and second vanishing points to restore parallel lines in the text image.

3. The method of claim 2 , wherein the plurality of image lines includes a plurality of vertical linelets and a plurality of horizontal linelets, each vertical linelet being a skeletal line through an upright portion of a connected component, each horizontal linelet being a regression line through a center of a set of adjacent connected components.

4. The method of claim 3 , wherein calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines further comprises:

calculating the horizontal vanishing point using the horizontal linelets; and

calculating the vertical vanishing point using the vertical linelets.

5. The method of claim 1 , wherein combining the first text image and the second text image includes blending the first text image with the second text image.

6. The method of claim 1 , wherein a number of matching shape descriptors between the first set of vectors associated with the first text image and the second set of vectors associated with the second text image satisfies a threshold number of matching shape descriptors.

7. The method of claim 1 , wherein combining the plurality of text images further comprises:

blending the plurality of text images that form the single text document;

de-skewing the single text document; and

performing optical character recognition on the single text document.

8. A system, comprising:

one or more computers;

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving a scan request, wherein the scan request includes a plurality of text images, each text image representing a portion of a text document, wherein the plurality of text images includes a first text image and a second text image that at least partially overlap;

for each text image of the plurality of text images:

rectifying the text image to generate a text image with parallel image lines,

generating a plurality of word bounding boxes that enclose one or more connected components in the text image, wherein each word bounding box is associated with a respective word of the text image, and

generating, for each respective word bounding box in the text image, a vector that describes a shape of the respective word;

combining text images of the plurality of text images that at least partially overlap to form a single text document including combining the first text image of the plurality of text images and the second text image of the plurality of text images by matching one or more shape descriptors from a first set of vectors generated for each respective word in the first text image and one or more shape descriptors from a second set of vectors generated for each respective word in the second text image; and

providing the combined image as a scannable document.

9. The system of claim 8 , wherein rectifying each text image of the plurality of text images comprises:

determining a plurality of connected components in the text image, each connected component being a filled portion of a symbol;

generating a plurality of image lines in the text image including vertical linelets and horizontal linelets based on the plurality of connected components;

calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines; and

applying a geometric formula to the first and second vanishing points to restore parallel lines in the text image.

10. The system of claim 9 , wherein the plurality of image lines includes a plurality of vertical linelets and a plurality of horizontal linelets, each vertical linelet being a skeletal line through an upright portion of a connected component, each horizontal linelet being a regression line through a center of a set of adjacent connected components.

11. The system of claim 10 , wherein calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines further comprises:

calculating the horizontal vanishing point using the horizontal linelets; and

calculating the vertical vanishing point using the vertical linelets.

12. The system of claim 8 , wherein combining the first text image and the second text image includes blending the first text image with the second text image.

13. The method of claim 8 , wherein a number of matching shape descriptors between the first set of vectors associated with the first text image and the second set of vectors associated with the second text image satisfies a threshold number of matching shape descriptors.

14. The system of claim 8 , wherein combining the plurality of text images further comprises:

blending the plurality of text images that form the single text document;

de-skewing the single text document; and

performing optical character recognition on the single text document.

15. A computer storage medium encoded with a computer program, the program comprising instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a scan request, wherein the scan request includes a plurality of text images, each text image representing a portion of a text document, wherein the plurality of text images includes a first text image and a second text image that at least partially overlap;

for each text image of the plurality of text images:

rectifying the text image to generate a text image with parallel image lines,

generating a plurality of word bounding boxes that enclose one or more connected components in the text image, wherein each word bounding box is associated with a respective word of the text image, and

generating, for each respective word bounding box in the text image, a vector that describes a shape of the respective word;

combining text images of the plurality of text images that at least partially overlap to form a single text document including combining the first text image of the plurality of text images and the second text image of the plurality of text images by matching one or more shape descriptors from a first set of vectors generated for each respective word in the first text image and one or more shape descriptors from a second set of vectors generated for each respective word in the second text image; and

providing the combined image as a scannable document.

16. The computer storage medium of claim 15 , wherein rectifying each text image of the plurality of text images comprises:

determining a plurality of connected components in the text image, each connected component being a filled portion of a symbol;

generating a plurality of image lines in the text image including vertical linelets and horizontal linelets based on the plurality of connected components;

calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines; and

applying a geometric formula to the first and second vanishing points to restore parallel lines in the text image.

17. The computer storage medium of claim 16 , wherein the plurality of image lines includes a plurality of vertical linelets and a plurality of horizontal linelets, each vertical linelet being a skeletal line through an upright portion of a connected component, each horizontal linelet being a regression line through a center of a set of adjacent connected components.

18. The computer storage medium of claim 17 , wherein calculating a horizontal vanishing point and a vertical vanishing point based on the plurality of image lines further comprises:

calculating the horizontal vanishing point using the horizontal linelets; and

calculating the vertical vanishing point using the vertical linelets.

19. The computer storage medium of claim 15 , wherein combining the first text image and the second text image includes blending the first text image with the second text image.

20. The method of claim 15 , wherein a number of matching shape descriptors between the first set of vectors associated with the first text image and the second set of vectors associated with the second text image satisfies a threshold number of matching shape descriptors.

21. The computer storage medium of claim 15 , wherein combining the plurality of text images further comprises:

blending the plurality of text images that form the single text document;

de-skewing the single text document; and

performing optical character recognition on the single text document.

22. The method of claim 1 , wherein generating a particular vector for a particular word includes generating a grid for the corresponding word bounding box, performing histogram oriented gradient over the cells of the grid to determine word shape descriptors, and generating the particular vector from the words shape descriptors.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2015
From: CHAUDHURY, KRISHNENDU; CHEN, LU; PETROU, DAVID; AGUERA-ARCAS, BLAISE
To: GOOGLE INC.
Reel/Frame 036388/0580 →
Continuity (1)
Related Publication 20160307059A1 · Oct 20, 2016