IP Library Granted Patent US 7,639,875
Granted Patent B2
US 7,639,875 · App. 12/245,447 · Granted Dec 29, 2009

System and method for capturing and processing business data

Assignee: ScanR, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,639,875
App. No.
12/245,447
Granted
Dec 29, 2009
Kind
B2
Abstract

A method and a system for interpreting information in a document are provided, with the system receiving an image of a document from a remote source and converting it into multiple sets of blocks of characters. Tags indicating likely meaning of blocks are assigned to them. At least some of the blocks have an associated score representing the probability that the characters in the block correctly represent the characters in the original image. The system selects one set from multiple sets based on the scores associated to certain blocks determined by accessing remote information over the Internet.

Claims (45)

1. A server device for use in interpreting information in a document, comprising:

a storage component arranged to receive and store an image of a document received from a remote source; and

a processor that includes data and instructions configured to perform actions, including:

representing the image as text that includes a plurality of characters, some of the characters in the plurality having alternative versions with associated confidence probabilities;

generating a set of tokenization's, each tokenization comprising a set of unique tokens that comprise collections of characters, wherein different tokens are defined for different versions of a character, and wherein for characters with different versions a single version is included in a tokenization;

assigning one or more tags to the tokens, the tags indicating a possible meaning of a corresponding token, and at least some of the tags having a score value indicating a probability of accuracy;

parsing each tokenization in the set of tokenizations based on a determined grammar to obtain multiple tokenizations with a single tag being assigned to each token;

assigning each tokenization an aggregate score based at least on compliance with the determined grammar; and

selecting as a final tokenization one tokenization with tags based on the aggregate score from the multiple tokenizations.

2. The server device of claim 1 , wherein information represented in the document comprises identifiable semantic structures and wherein a subset of the semantic structures being presented only once and having a unique meaning.

3. The server device of claim 1 , wherein assigning tags further comprises filtering tokens of the tokenizations by identifying selected tokens as common words in the document and assigning tags to neighboring words based on a position relative to the common words.

4. The server device of claim 1 , wherein parsing each tokenization is configured to begin anywhere in the document and includes searching both forward and backward to satisfy conditions of grammar rules.

5. The server device of claim 1 , wherein the processor is configured to perform actions, further including:

converting the final tokenization into a data structure wherein the tags specify located fields of the data structure and the tokens provide data for such fields.

6. The server device of claim 1 , wherein selecting as a final tokenization further comprises:

providing a first portion of a given tokenization from the multiple tokenizations to an external database to find a record matching the first portion.

7. The server device of claim 6 , wherein selecting as a final tokenization further comprises:

determining whether a match exists between a second portion of the given tokenization with information in the record; and

increasing the score of the given tokenization as the final tokenization if the match is detected.

8. A computer-readable storage medium that includes data and instructions, wherein the execution of the instructions on a server device provides for interpreting information in a document by enabling actions, comprising:

receiving an image of the document over a network from a remote source;

converting the image into multiple sets of blocks of characters, each block having tags indicating an associated meaning and at least some of the blocks having an associated score representing a probability that the characters in the block correctly represent the image;

parsing each set of blocks based on a predetermined grammar to remove certain tags, leaving a single tag per block; and

selecting a final set from the multiple sets based on the scores associated with at least some of blocks and based on information provided as a result of accessing remote information over the network.

9. The computer-readable storage medium of claim 8 , wherein execution of the instructions enable actions, further comprising:

providing the final set to at least one of a mail account, or another storage medium.

10. The computer-readable storage medium of claim 8 , wherein the parsing software is configured to parse anywhere in the document and further includes searching both forward and backward to satisfy conditions of grammar rules.

11. The computer-readable storage medium of claim 8 , wherein the remote device from which the image is received includes an image capture capability.

12. The computer-readable storage medium of claim 8 , wherein selecting the final set further comprises providing a first one or more blocks of a given set from the multiple sets to an external database to locate a record matching the one or more blocks.

13. The computer-readable storage medium of claim 8 , wherein selecting the final set further comprises searching websites over the network using a first one or more blocks of a given set from the multiple sets to locate content matching the first one or more blocks.

14. The computer-readable storage medium of claim 8 , wherein converting the image further comprises forming sets of groups of characters and assigning one or more tags to each set.

15. The computer-readable storage medium of claim 8 , wherein the final set is converted into a data structure with tags specifying located fields in the data structure.

16. The computer-readable storage medium of claim 8 , wherein information represented in the document comprises identifiable semantic structures, and wherein a subset of the semantic structures has a unique meaning.

17. A system that is configured to interpret information in a document, comprising:

a receiving component configured to receive an image of the document over a network; and

a processor executing instructions on a computer that perform actions, comprising:

converting the image into multiple sets of blocks of characters, each block having tags indicating an associated meaning and at least some of the blocks having an associated score representing a probability that the characters in the block correctly represent the image;

parsing each set of blocks based on a predetermined grammar to remove certain tags, leaving a single tag per block; and

selecting a final set from the multiple sets based on the scores associated with at least some of blocks, and based on information provided as a result of accessing remote content over the network.

18. The system of claim 17 , wherein selecting the final set further comprises:

searching for content using a first one or more blocks of a given set from the multiple sets to find content matching the first one or more blocks; and

if a match between a second one or more blocks of the given set with at least some of the content is detected, increasing a score of the given set as the final set based on an amount of the content that matches.

19. The system of claim 17 , wherein the system is configured to operate as one of a server, a client, or a mobile device.

20. The system of claim 17 , wherein the processor executes instructions that perform actions, further comprising:

assigning tags to blocks by looking up the blocks in a dictionary, wherein the dictionary comprises tags and scores representing a probability of the tag assignment being correct.

Assignments (5)
SECURITY INTEREST Recorded Dec 18, 2014
From: FLINT MOBILE, INC.
To: VENTURE LENDING & LEASING VII, INC.
Reel/Frame 034549/0432 →
SECURITY AGREEMENT Recorded Nov 29, 2012
From: FLINT MOBILE, INC.
To: VENTURE LENDING & LEASING VI, INC.
Reel/Frame 029380/0271 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2012
From: SCANR, INC.
To: FLINT MOBILE, INC.
Reel/Frame 028008/0678 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2009
From: MOLNAR, JOSEPH; FERREIRA, PAULO; NIEUWLAND, DAN; MUTZ, ANDREW H.
To: MOBILESCAN
Reel/Frame 022840/0419 →
CHANGE OF NAME Recorded Jun 17, 2009
From: MOBILESCAN, INC.
To: SCANR, INC.
Reel/Frame 022840/0501 →
Continuity (3)
Continuation 1117659200 · Jul 6, 2005
Continuation In Part 1113304900 · May 18, 2005
Related Publication 20090034844A1 · Feb 5, 2009