IP Library Granted Patent US 10,769,425
Granted Patent B2
US 10,769,425 · App. 16/101,763 · Granted Sep 8, 2020

Method and system for extracting information from an image of a filled form document

Inventors: Antonio Foncubierta Rodriguez (Zurich, CH); Maria Gabrani (Thalwil, CH); Guillaume Jaume (Zurich, CH)
Assignee: International Business Machines Corporation
G06K9/00449G06K9/00456G06K9/00463G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,425
App. No.
16/101,763
Granted
Sep 8, 2020
Kind
B2
Abstract

A method of determining a hierarchy of a blank template using an image of the blank template and using the determined hierarchy for providing labels and field values of text lines of a filled form document.

Claims (91)

1. A method of extracting information from an image of a filled form document, the form document comprising labeled fields comprising values and arranged in sections of the document, the method comprising:

providing an image of a blank template of the form document;

extracting first text lines from the image of the blank template using a text line recognizer of an optical character recognition (OCR) system;

performing using the OCR system an optical character recognition of the extracted first text lines resulting in first machine-encoded text lines;

merging the first machine encoded text lines into candidate sections based on the location of the first text lines in the blank template;

evaluating for each candidate section a set of predefined features indicative of the text lines and their location;

combining the evaluated features of each candidate section to generate an identifier of the candidate section, resulting in multiple identifiers;

determining sets of candidate sections that each share a respective identifier of the multiple identifiers;

selecting a set of candidate sections of the sets of sections that fulfills a predefined selection criterion based on the number of sections and number of respective fields;

determining as the hierarchy of the blank template the selected set of candidate sections;

extracting second text lines from the image of the filled form document using the text line recognizer;

performing using the OCR system an optical character recognition of the extracted second text lines resulting in second machine-encoded text lines;

merging the second machine encoded text lines into text regions based on the location of the second text lines;

for each second text line of the filled form document:

identifying the section of the determined hierarchy that matches the second text line based on the region of the second text line;

determining the first text line of the section that corresponds to the second text line by comparing the second text line with the first text lines, and determining the label of the field of the second text line and the position of the label within the filled form using the determined first text line;

using the determined labels for identifying and extracting field values of the second text lines;

comparing each of the field values with the determined labels for assigning the field value to a respective label of the determined labels.

2. The method of claim 1 , wherein comparing the first and second text lines is performed using a similarity metric indicative of the similarity between words of the first and second text lines.

3. The method of claim 2 , the similarity metric comprises at least one of Levenshtein distance, word accuracy (WA), combined word accuracy (CWA) and field detection accuracy (FDA).

4. The method of claim 1 , wherein the selection criterion is the following:

min

h

S

1

S

i

=

0

S

-

1

Q

i

,

where S is the number of candidate sections shared by a given identifier and Q is the number of fields in a given candidate section.

5. The method of claim 1 , wherein the selection criterion is the following:

min

h

S

median

{

Q

}

,

where S is the number of candidate sections shared by a given identifier and median{Q} is the median number of fields for a given hierarchy.

6. The method of claim 1 , wherein the combining the evaluated features into an identifier h is performed as follows h=Σ n=0 M f n ·2 n where f n represents the n th feature of M features.

7. The method of claim 1 , wherein the set of predefined features comprises at least one of: font size, font type, position of a character, background color, indication of a character being an uppercase or lowercase character.

8. The method of claim 1 , wherein the comparing of the field value with a label of the determined labels comprises computing a probability that the field value and the label corresponds to each other using visual features, wherein the visual features comprise L1 distance between centroids of bounding boxes of the compared label and field value, L1 distance between top-left corners of bounding boxes of the compared label and field value, L1 distance between closest borders of bounding boxes of the compared label and field value.

9. The method of claim 1 , wherein the merging of the machine encoded text lines is performed using a vertical merging technique.

10. The method of claim 1 , wherein the identifier of the text region comprises a hash value.

11. The method of claim 1 , further comprising providing the hierarchy of the filled form document as an xml file comprising relative positions of the text lines and sections.

12. The method of claim 1 , wherein the field value comprises a value that is input of the field and/or information indicative of the field.

13. A computer program product comprising a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code configured to implement steps of the method according to claim 1 .

14. A computer system for extracting information from an image of a filled form document, the form document comprising labeled fields comprising values and arranged in sections of the document, the computer system comprising an image of a blank template of the form document, the computer system being configured for:

extracting first text lines from the image of the blank template using a text line recognizer of an optical character recognition (OCR) system;

performing using the OCR system an optical character recognition of the extracted first text lines resulting in first machine-encoded text lines;

merging the first machine encoded text lines into candidate sections based on the location of the first text lines in the blank template;

evaluating for each candidate section a set of predefined features indicative of the text lines and their location;

combining the evaluated features of each candidate section to generate an identifier of the candidate section, resulting in multiple identifiers;

determining sets of candidate sections that each share a respective identifier of the multiple identifiers;

selecting a set of candidate sections of the sets of sections that fulfills a predefined selection criterion based on the number of sections and number of respective fields;

determining as the hierarchy of the blank template the selected set of candidate sections;

extracting second text lines from the image of the filled form document using the text line recognizer;

performing using the OCR system an optical character recognition of the extracted second text lines resulting in second machine-encoded text lines;

merging the second machine encoded text lines into text regions based on the location of the second text lines;

for each second text line of the filled form document:

identifying the section of the determined hierarchy that matches the second text line based on the region of the second text line;

determining the first text line of the section that corresponds to the second text line by comparing the second text line with the first text lines, and determining the label of the field of the second text line and the position of the label within the filled form using the determined first text line;

using the determined labels for identifying and extracting field values of the second text lines;

comparing each of the field values with the determined labels for assigning the field value to a respective label of the determined labels.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2018
From: GABRANI, MARIA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046649/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2018
From: FONCUBIERTA RODRIGUEZ, ANTONIO; JAUME, GUILLAUME
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046628/0291 →
Continuity (1)
Related Publication 20200050845A1 · Feb 13, 2020
Cited By (1)
US 12,494,076