IP Library Granted Patent US 10,943,105
Granted Patent B2
US 10,943,105 · App. 16/580,719 · Granted Mar 9, 2021

Document field detection and parsing

Inventors: Shuo Chen (Livingston, NJ); Venkataraman Pranatharthiharan (Exton, PA)
Assignee: The Neat Company, Inc.
G06K9/00442G06K9/00449G06K9/46G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,943,105
App. No.
16/580,719
Granted
Mar 9, 2021
Kind
B2
Abstract

A system and method for invoice field detection and parsing includes the steps of extracting character bounding blocks using optical character recognition (OCR) or digital character extraction (DCE), enhancing the image quality, analyzing the document layout based on imaging techniques, detecting the invoice field based on the machine learning techniques, and parsing the invoice field value based on the content information.

Claims (46)

1. A method of detecting and extracting at least one of character and number values in a document field of a received analog or digital financial document, comprising:

providing the analog or digital financial document as an input image or a PDF input;

extracting an image layer, when the analog or digital financial document is the PDF input;

extracting text and character blocks from the input image or the extracted image layer;

extracting connected components from the input image or the extracted image layer and analyzing the extracted connected components to determine document layout information identifying a structure of the input image or the extracted image layer, the document layout information including character information and noncharacter information including at least one of barcodes, tables, and logos in said input image or the extracted image layer;

detecting a document field from the document layout information of the input image or the extracted image layer; and

parsing the detected document field to extract at least one of character and number values from the extracted text and character blocks of the detected document field of the input image or the extracted image layer.

2. A method as in claim 1 , wherein extracting text and character blocks comprises processing the input image or the extracted image layer using optical character recognition (OCR).

3. A method as in claim 1 , further comprising enhancing an image quality of the input image or the extracted image layer prior to the analysis to determine the document layout information.

4. The method of claim 3 , wherein enhancing the image quality of the input image or the extracted image layer comprises:

rotating and deskewing the input image or the extracted image layer;

cropping the rotated and deskewed input image or the extracted image layer; and

enhancing the background of the input image or the extracted image layer.

5. The method of claim 4 , wherein enhancing the background of the input image or the extracted image layer comprises computing a grayscale histogram of the input image or the extracted image layer, searching for a grayscale value with maximum count in said histogram and treating the grayscale value as a median grayscale value of the background.

6. The method of claim 4 , wherein enhancing the background of the input image or the extracted image layer comprises generating a grayscale interval originated from said median grayscale value.

7. The method of claim 4 , wherein enhancing the background of the input image or the extracted image layer comprises resetting the grayscale value of background pixels of the background if the grayscale value of the background pixels falls outside of the gray scale interval.

8. The method of claim 1 , wherein analyzing the input image or the extracted image layer to determine document layout information comprises creating a tree data structure to store the document layout information.

9. The method of claim 1 , wherein analyzing the extracted connected components to determine document layout information comprises:

converting the input image or the extracted image layer to a binary image;

extracting connected components from the binary image;

removing noise from the connect components to generated clean connected components;

detecting at least one of barcodes, tables, logos, and characters in said clean connected components;

generating text lines from said characters; and

generating words and paragraph zones from said text lines.

10. The method of claim 9 , wherein detecting characters in said connected components comprises collecting all components that are not noise, barcodes, logos, nor table grid lines.

11. The method of claim 9 , wherein generating text lines from said characters comprises sorting said characters by their left coordinates.

12. The method of claim 9 , wherein generating text lines from said characters comprises identifying whether a character belongs to a text line by:

comparing a height of said character and said text line and determining whether said character belongs to said text line based on a height difference;

computing a vertical overlapping between said character and said text line and determining whether said character belongs to said text line based on an overlapping length; and/or

comparing a horizontal distance between said character and said text line and determining whether said character belongs to said text line based on the horizontal distance.

13. The method of claim 9 , wherein generating text lines from said characters comprises combining overlapped text line candidates.

14. The method of claim 9 , wherein generating text lines from said characters comprises removing a text line thinner than a threshold from a text line candidate.

15. The method of claim 9 , wherein generating words from said text lines comprises estimating a word space from a text line.

16. The method of claim 15 , wherein generating words from said text lines comprises identifying whether a character belongs to a word by:

comparing a height of said character and said word and determining whether said character belongs to said word based on a height difference;

comparing a horizontal distance between said character and said word with said estimated word space and determining whether said character belongs to said word;

comparing a height of said character and a character immediately after said character and determining whether said character belongs to said word based on a height difference; and/or

comparing a horizontal distance between said character and a character immediately after said character with said estimated word space and determining whether said character belongs to said word.

17. The method of claim 1 , wherein detecting the document field from the document layout information comprises running a text line through a cascade classifier until said text line is matched with one document field or is discarded.

18. The method of claim 1 , wherein parsing the detected document field to extract at least one of character and number values from the input image or the extracted image layer comprises applying to the input image or the extracted image layer a series of independent content-based field parsers, each of which corresponds to a document field and parses out a corresponding field value.

19. The method of claim 1 , wherein extracting the image layer from the PDF input comprises:

checking a document format to determine when the financial document is the PDF input;

when the financial document is the PDF input, extracting the image layer from the PDF input;

checking a text layer of the PDF input to determine whether the PDF input is searchable or non-searchable; and

extracting the text and character blocks from the PDF input when the PDF input is searchable.

20. The method of claim 19 , wherein checking the document format of the PDF input comprises extracting a document extension from a document name and determining the document format from the document extension.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Jan 9, 2023
From: PACIFIC WESTERN BANK
To: THE NEAT COMPANY, INC.
Reel/Frame 062317/0169 →
SECURITY INTEREST Recorded Mar 9, 2021
From: THE NEAT COMPANY, INC.
To: PACIFIC WESTERN BANK
Reel/Frame 055536/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2019
From: CHEN, SHUO; PRANATHARTHIHARAN, VENKATARAMAN
To: THE NEAT COMPANY, INC. D/B/A NEATRECEIPTS, INC.
Reel/Frame 050480/0831 →
Cited By (1)
US 12,423,465