IP Library Granted Patent US 9,008,447
Granted Patent B2
US 9,008,447 · App. 11/547,118 · Granted Apr 14, 2015

Method and system for character recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,008,447
App. No.
11/547,118
Granted
Apr 14, 2015
Kind
B2
Abstract

A method and system for character recognition are described. In one embodiment, it may use matched sequences rather than character shape to determine a computer legible result.

Claims (67)

1. A method of character recognition comprising:

capturing, by a character-recognition system, a scan of a document;

segmenting, by the character-recognition system, the scan into a sequence of character images;

generating, by the character-recognition system, a sequence of identifiers based on the character images, wherein each identifier in the sequence of identifiers represents a particular character, and wherein the generating comprises:

assigning an identifier to a first character image in the sequence of images;

for each subsequent character image in the sequence of images, comparing the subsequent character image to one or more character images that precede the subsequent character image in the sequence of character images;

for each subsequent character image that matches one of the previous character images based on the comparison, assigning the identifier that is assigned to the one character image to the subsequent image; and

for each subsequent character image that does not match one of the previous character images based on the comparison, assigning a new identifier that is different from each identifier previously assigned; and

obtaining, by the character-recognition system, text based on the sequence of identifiers, the obtaining comprising:

comparing the sequence of identifiers to a plurality of identifier sequences, each identifier sequence corresponding to a sequence of characters; and

identifying as the text a sequence of characters corresponding to an identifier sequence that matches the sequence of identifiers.

2. The method of claim 1 , wherein the text comprises text corresponding to the one or more character images.

3. The method of claim 1 , wherein generating the sequence of identifiers based on the one or more character images comprises:

initializing a value of an identifier; and

for each character image of the character images:

determining a character represented by the character image;

determining whether an identifier representing the character is already stored as part of the sequence of identifiers;

in response to determining that the identifier representing the character is not already stored as part of the sequence of identifiers:

incrementing the value of the identifier;

storing the value of the identifier as part of the sequence of identifiers; and

outputting the value of the identifier.

4. The method of claim 1 , wherein obtaining the text comprises querying a database.

5. The method of claim 1 , wherein the text comprises one or more words.

6. The method of claim 5 , further comprising:

determining a probability of a combination of the one or more words.

7. The method of claim 6 , further comprising:

determining whether the probability is less than a probability threshold; and

in response to determining that the probability is less than the probability threshold, determining that the text is in error.

8. The method of claim 1 , wherein the scan comprises an image of a portion of a document.

9. The method of claim 1 , wherein the scan comprises an image of one or more words.

10. The method of claim 9 , wherein the scan comprises an image of fewer than 20 words.

11. The method of claim 1 , wherein the text consists of characters or character representations.

12. The method of claim 1 , wherein at least one character image of the character images comprises an image of at least two characters.

13. The method of claim 1 , further comprising:

generating an output stream comprising one or more offsets of the one or more character images.

14. The method of claim 13 , wherein a first offset of the one or more offsets is based on first and second images of the character images,

wherein the first image and the second image both represent a first character, and

wherein the first offset is based on a number of intervening images of the character images between the first image and second image.

15. A character-recognition system, comprising:

memory, configured to store at least a scan of a document; and

a processor, configured to:

segment the scan into character images;

generate a sequence of identifiers based on the character images, wherein each identifier in the sequence of identifiers represents a particular character, and wherein the generating comprises:

assigning an identifier to a first character image in the sequence of images;

for each subsequent character image in the sequence of images, comparing the subsequent character image to one or more character images that precede the subsequent character image in the sequence of character images;

for each subsequent character image that matches one of the previous character images based on the comparison, assigning the identifier that is assigned to the one character image to the subsequent image; and

for each subsequent character image that does not match one of the previous character images based on the comparison, assigning a new identifier that is different from each identifier previously assigned; and

obtain text based on the sequence of identifiers by:

comparing the sequence of identifiers to a plurality of identifier sequences, each identifier sequence corresponding to a sequence of characters; and

identifying as the text a sequence of characters corresponding to an identifier sequence that matches the sequence of identifiers.

16. The character-recognition system of claim 15 , further comprising a power source configured to power the system.

17. The character-recognition system of claim 15 , further comprising an input comprising an optical sensor configured to: generate the scan and store the scan in the memory.

18. The character-recognition system of claim 15 , wherein the memory is further configured to store at least the text.

19. A system, comprising:

means for capturing a scan of a document;

means for segmenting the scan into a sequence of character images;

means for generating a sequence of identifiers based on the character images, wherein each identifier in the sequence of identifiers represents a particular character, and wherein the generating comprises:

assigning an identifier to a first character image in the sequence of images;

for each subsequent character image in the sequence of images, comparing the subsequent character image to one or more character images that precede the subsequent character image in the sequence of character images;

for each subsequent character image that matches one of the previous character images based on the comparison, assigning the identifier that is assigned to the one character image to the subsequent image; and

for each subsequent character image that does not match one of the previous character images based on the comparison, assigning a new identifier that is different from each identifier previously assigned; and

means for obtaining text based on the sequence of identifiers, the obtaining comprising:

comparing the sequence of identifiers to a plurality of identifier sequences, each identifier sequence corresponding to a sequence of characters; and

identifying as the text a sequence of characters corresponding to an identifier sequence that matches the sequence of identifiers.

20. The system of claim 19 , further comprising means for printing the document.

21. The method of claim 1 , wherein comparing the subsequent character image to the one or more character images that precede the subsequent character image comprises subtracting each of the one or more previous character images from the subsequent character image.

22. The method of claim 21 , further comprising determining that the one character image matches the subsequent character image based on a resulting pattern that results from subtracting the one character image from the subsequent character image satisfies a size threshold.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Aug 3, 2021
From: GOOGLE LLC
To: KYOCERA CORPORATION
Reel/Frame 057651/0445 →
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044334/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2011
From: EXBIBLIO B.V.
To: GOOGLE INC.
Reel/Frame 025779/0251 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2007
From: KING, MARTIN T.; GROVER, DALE L.; KUSHLER, CLIFFORD A.; STAFFORD-FRASER, JAMES Q.
To: EXBIBLIO B.V.
Reel/Frame 019799/0296 →