IP Library Granted Patent US 9,767,354
Granted Patent B2
US 9,767,354 · App. 15/146,848 · Granted Sep 19, 2017

Global geographic information retrieval, validation, and normalization

Inventors: Stephen Michael Thompson (Oceanside, CA); Jan W. Amtrup (Silver Spring, MD); Anthony Macciola (Irvine, CA)
Assignee: KOFAX, INC.
G06K9/00469G06F17/3028G06F17/30256G06F17/30371G06F17/30542G06K9/00442G06K9/00456G06K9/03G06K9/18G06K9/481G06Q10/10H04N1/40H04N1/40062
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,767,354
App. No.
15/146,848
Granted
Sep 19, 2017
Kind
B2
Abstract

According to one embodiment, a computer-implemented method includes: capturing an image of a document using a camera of a mobile device; performing optical character recognition (OCR) on the image of the document; extracting an identifier of the document from the image based at least in part on the OCR; comparing the identifier with content from one or more reference data sources, wherein the content from the one or more reference data sources comprises global address information; and determining whether the identifier is valid based at least in part on the comparison. The method may optionally include normalizing the extracted identifier, retrieving additional geographic information, correcting OCR errors, etc. based on comparing extracted information with reference content. Corresponding systems and computer program products are also disclosed.

Claims (57)

1. A computer-implemented method, comprising:

capturing an image of a document using a camera of a mobile device;

performing optical character recognition (OCR) on the image of the document;

extracting an identifier of the document from the image based at least in part on the OCR;

comparing the identifier with content from one or more reference data sources, wherein the content from the one or more reference data sources comprises global address information; and wherein the content from the one or more reference data sources is derived from geographic information organized in one or more of a proprietary address database and an open source address database; and wherein deriving the content from the geographic information comprises:

obtaining the geographic information from one or more of the proprietary address database and an open source address database; and

parsing the geographic information according to a set of predefined heuristic rules, wherein the set of predefined heuristic rules are configured to normalize the global address information obtained from the one or more sources according to a single convention for representing address information; and

determining whether the identifier is valid based at least in part on the comparison.

2. The method as recited in claim 1 , wherein the identifier consists of characters selected from a predefined alphabet, wherein the predefined alphabet consists of one or more of numerals, alphabetic characters, and symbols.

3. The method as recited in claim 1 , wherein the identifier comprises a partial or complete address.

4. The method as recited in claim 1 , wherein the identifier comprises one or more of:

a street name, a street number, a block number, a unit number, a city name, a county name, a municipality name, a state name, a state abbreviation, a country name, a country abbreviation, and a ZIP code.

5. The method as recited in claim 1 , the comparing comprising fuzzy matching the identifier with the content from the one or more data sources.

6. The method as recited in claim 1 , wherein the identifier is validated based at least in part on determining a fuzzy match exists between the identifier and at least a portion of the global address information, wherein the fuzzy match is characterized by no more than two character mismatches between the identifier and at least the portion of the global address information.

7. The method as recited in claim 1 , comprising locating the identifier within the image based on a connected components analysis.

8. The method as recited in claim 1 , wherein the OCR is performed only on a portion of the image determined to depict the identifier.

9. The method as recited in claim 1 , comprising determining a locality associated with the identifier; and

wherein the set of predefined heuristic rules are selected based on the locality determined to be associated with the extracted identifier.

10. The method as recited in claim 1 , wherein deriving the content from the geographic information comprises populating the one or more data sources with the content, wherein the content consists of geographic information parsed using the set of predefined heuristic rules.

11. The method as recited in claim 1 , wherein deriving the content from the geographic information comprises normalizing the geographic information to expand one or more abbreviations present in the geographic information; and

wherein the content excludes abbreviated geographic information.

12. The method as recited in claim 1 , comprising normalizing the extracted identifier prior to comparing the identifier with content from one or more reference data sources, wherein the normalizing is performed according to one or more predefined business rules corresponding to a particular locality.

13. The method as recited in claim 1 , comprising determining a locality corresponding to the extracted identifier, and

retrieving additional geographic information associated with a location corresponding to the identifier based at least in part on the locality.

14. The method as recited in claim 13 , wherein determining the locality is based at least in part on one or more of a content and a format of the identifier.

15. The method as recited in claim 13 , wherein retrieving the additional geographic information is based at least in part on latitude and longitude coordinates corresponding to the identifier.

16. The method as recited in claim 1 , comprising at least one of:

detecting one or more OCR errors based at least in part on textual information from a complementary document;

detecting one or more OCR errors using one or more predefined business rules;

detecting one or more OCR errors based at least in part on textual information from the complementary document and one or more of the predefined business rules;

correcting at least one detected OCR error using one or more of the predefined business rules;

correcting at least one detected OCR error using textual information from the complementary document;

correcting at least one detected OCR error using textual information from the complementary document and one or more of the predefined business rules;

normalizing data from a complementary document using at least one of the predefined business rules;

normalizing data from the document using at least one of the predefined business rules; and

normalizing data from the document using textual information from the complementary document and at least one of the predefined business rules.

17. A computer program product, comprising a non-transitory computer readable storage medium having stored/encoded thereon computer readable program instructions configured to cause a processor, upon execution thereof, to:

receive an image of a document;

perform optical character recognition (OCR) on the image of the document;

extract an identifier of the document from the image based at least in part on the OCR;

compare the identifier with content from one or more reference data sources, wherein the content from the one or more reference data sources comprises global address information; and wherein the content from the one or more reference data sources is derived from geographic information organized in one or more of a proprietary address database and an open source address database; and wherein deriving the content from the geographic information comprises:

obtaining the geographic information from one or more of the proprietary address database and an open source address database; and

parsing the geographic information according to a set of predefined heuristic rules, wherein the set of predefined heuristic rules are configured to normalize the global address information obtained from the one or more sources according to a single convention for representing address information; and

determine whether the identifier is valid based at least in part on the comparison.

18. A computer-implemented method, comprising:

capturing an image using a camera of a mobile device;

classifying the image as an image of a document, wherein the classifying comprises:

generating a first feature vector representative of the document, based on analyzing the image; and

comparing the first feature vector to a plurality of reference feature matrices;

performing optical character recognition (OCR) on the image of the document;

extracting an identifier of the document from the image based at least in part on the OCR;

comparing the identifier with content from one or more reference data sources, wherein the content from the one or more reference data sources comprises global address information; and wherein the content from the one or more reference data sources is derived from geographic information organized in one or more of a proprietary address database and an open source address database; and wherein deriving the content from the geographic information comprises:

obtaining the geographic information from one or more of the proprietary address database and an open source address database; and

parsing the geographic information according to a set of predefined heuristic rules, wherein the set of predefined heuristic rules are configured to normalize the global address information obtained from the one or more sources according to a single convention for representing address information;

determining whether the identifier is valid based at least in part on the comparison;

associating the image of the document with metadata descriptive of one or more of the document and information relating to the document; and

storing the image of the document and the associated metadata to a memory of the mobile device.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2024
From: KOFAX, INC.
To: TUNGSTEN AUTOMATION CORPORATION
Reel/Frame 067428/0392 →
RELEASE OF SECURITY INTEREST Recorded Jul 21, 2022
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: KAPOW TECHNOLOGIES, INC.; KOFAX, INC.
Reel/Frame 060805/0161 →
FIRST LIEN INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 20, 2022
From: KOFAX, INC.; PSIGEN SOFTWARE, INC.
To: JPMORGAN CHASE BANK, N.A. AS COLLATERAL AGENT
Reel/Frame 060757/0565 →
SECURITY INTEREST Recorded Jul 20, 2022
From: KOFAX, INC.; PSIGEN SOFTWARE, INC.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 060768/0159 →
SECURITY INTEREST Recorded Jul 7, 2017
From: KOFAX, INC.
To: CREDIT SUISSE
Reel/Frame 043108/0207 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2016
From: THOMPSON, STEPHEN MICHAEL; AMTRUP, JAN W.; MACCIOLA, ANTHONY
To: KOFAX, INC.
Reel/Frame 038736/0084 →
Continuity (6)
Continuation In Part 14588147 · Dec 31, 2014
Continuation 14176006 · Feb 7, 2014
Continuation In Part 13948046 · Jul 22, 2013
Continuation 13691610 · Nov 30, 2012
Continuation 12368685 · Feb 10, 2009
Related Publication 20160328610A1 · Nov 10, 2016