IP Library Granted Patent US 8,049,921
Granted Patent B2
US 8,049,921 · App. 12/082,040 · Granted Nov 1, 2011

System and method for transferring invoice data output of a print job source to an automated data processing system

Assignee: Bottomline Technologies (de) Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,049,921
App. No.
12/082,040
Granted
Nov 1, 2011
Kind
B2
Abstract

A data capture system receives a sequence of document objects and, for each, writes output data values to a structure. A first tier extraction system is adapted to receive each document object. For each required data element, the first tier extraction system obtains identification of a positional element value from a positional data set that includes, as its data element, identification of the required data element; and, if the document object includes a qualifying text string, writes an output data value to the output data structure in association with identification of the required data element. A second tier extraction system receives each such document object that does not include a qualifying text string, performs character recognition on a graphical representation thereof and, for each required data element, writes an output data value to the output data structure in association with identification of the required data element.

Claims (87)

1. A data capture system for receipt of a sequence of at least one output document object and, for each output document object, writing output data values to an output data structure, the system comprising:

a non-transitory data storage comprising:

a positional identification storage including at least two positional data sets, each positional data set includes:

i) identification of a required invoice data element; and

ii) identification of a positional element value defining a location within a graphical representation of each output document object at which a text string representative of a value of the required invoice data element is positioned;

a first tier data extraction system adapted to receive each output document object, each output document object being in a print language format comprising a plurality of print elements, each print element including a print component and at least one position identifier value identifying a position at which the print component is rendered within a graphical representation of the output document object, each print component being one of: i) a text string representing a value of an invoice data element; and ii) a graphic image, the first tier data extraction system being further adapted to, for each received output document object:

for each required invoice data element:

obtain the identification of the positional element value from the positional data set that includes, as its invoice data element, identification of the required invoice data element;

if the output document object includes a qualifying text string, write an output data value to the output data structure in association with identification of the required invoice data element, the output data value being one of: i) at least a portion of the qualifying text string; and ii) a numerical value represented by at least a portion of the qualifying text string, wherein a qualifying text string is a text string of a print element that includes a position identifier value that is within a predetermined variance from the positional element value of the positional data set; and

if the output document object does not include a qualifying text string, identifying the output document object for tier two processing;

a second tier data extraction system adapted to receive, for each output document object identified for tier two processing, a tier two document, the tier two document being the graphical representation of the output document object, the second tier data extraction system being further adapted to:

perform character recognition on the tier two document and construct a plurality of character recognition data sets, each character recognition data set associating a recognized character string within the tier two document and an identification of its location within the tier two document; and

for each required invoice data element for which the first tier data extraction system failed to write an output data value to the output data structure:

obtain the identification of the positional element value from the positional data set that includes, as its invoice data element, identification of the required invoice data element; and

if a character recognition data set includes a qualifying recognized character string, write an output data value to the output data structure in association with identification of the required invoice data element, the output data value being one of: i) at least a portion of the qualifying recognized character string; and ii) a numerical value represented by at least a portion of the qualifying recognized character string, wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value that is within a predetermined variance from the positional element value of the positional data set.

2. The data capture system of claim 1 , wherein the second tier data extraction system is further adapted to:

if a qualifying recognized character string is not included in any character recognition data set constructed for a tier two document, identifying the tier two document for tier three processing; and

a third tier identification system adapted to, for each tier two document identified for tier three processing:

generate a graphical representation of the tier two document at a workstation; and

for each required invoice data element for which the second tier data extraction system failed to write an output data value to the output data structure:

prompt for user input of an output data value;

receive user input of the output data value from the workstation; and

write, to the output data structure, the output data value received from the workstation in association with identification of the required invoice data element.

3. The data capture system of claim 2 , wherein:

at least one positional element value includes an abscissa value and an ordinate value defining a Cartesian coordinate within the graphical representation of the output document object at which an origin of the text string is positioned; and

a qualifying text string is a text string of a print element that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

4. The data capture system of claim 3 , wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

5. The data capture system of claim 2 , wherein:

at least one positional element value includes:

a reference to a second invoice data element;

an abscissa value; and

an ordinate value;

wherein: i) the abscissa value added to an abscissa value of the second invoice data element; and ii) the ordinate value added to an ordinate value of the second invoice data element define a Cartesian coordinate within the graphical representation of the output document object at which an origin of the text string is positioned; and

wherein a qualifying text string is a text string of a print element that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

6. The data capture system of claim 5 , wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

7. The data capture system of claim 2 , wherein:

the graphic image of at least one print component includes a pixelized representation of at least one character; and

the recognized character string of at least one character recognition data set includes characters matching characters of the pixelized representation of at least one character.

8. The data capture system of claim 2 , further comprising an accounting server, the accounting server:

crediting an account for a first charge in the event all required invoice data elements are written to the output data structure by the first tier data extraction system;

crediting the account for a second charge, different than the first charge, in the event:

any required invoice data elements are written to the output data structure by the second tier data extraction system; and

the output document object is not identified for tier three processing; and

crediting the account for a third charge, different than both the first charge and the second charge, in the event the output document object is identified for the tier three processing.

9. A method for capturing data from a sequence of at least one output document object and, for each output document object, writing output data values to an output data structure, the method comprising:

storing at least two positional data sets in a non-transitory data storage each positional data set includes:

i) identification of a required invoice data element; and

ii) identification of a positional element value defining a location within a graphical representation of each output document object at which a text string representative of a value of the required invoice data element is positioned;

receiving each output document object, each output document object being in a print language format comprising a plurality of print elements, each print element including a print component and at least one position identifier value identifying a position at which the print component is rendered within a graphical representation of the output document object, each print component being one of: i) a character string representing a value of an invoice data element; and ii) a graphic image;

for each required invoice data element, performing a first tier data extraction process, the first tier data extraction process comprising:

obtaining the identification of the positional element value from the positional data set that includes, as its invoice data element, identification of the required invoice data element;

if the output document object includes a qualifying text string, write an output data value to the output data structure in association with identification of the required invoice data element, the output data value being one of: i) at least a portion of the qualifying text string; and ii) a numerical value represented by at least a portion of the qualifying text string, wherein a qualifying text string is a text string of a print element that includes a position identifier value that is within a predetermined variance from the positional element value of the positional data set; and

if the output document object does not include a qualifying text string, identifying the output document object for tier two processing;

for each output document object identified for tier two processing, perform a second tier data extraction process, the second tier data extraction process comprising:

performing character recognition on a graphical representation of the output document object and construct a plurality of character recognition data sets, each character recognition data set associating a recognized character string within the graphical representation with an identification of its location within the graphical representation; and

for each required invoice data element for which an output data value was not written to the output data structure by the first tier data extraction process:

obtain the identification of the positional element value from the positional data set that includes, as its invoice data element, identification of the required invoice data element; and

if a character recognition data set includes a qualifying recognized character string, write an output data value to the output data structure in association with identification of the required invoice data element, the output data value being one of: i) at least a portion of the qualifying recognized character string; and ii) a numerical value represented by at least a portion of the qualifying recognized character string, wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value that is within a predetermined variance from the positional element value of the positional data set.

10. The method of claim 9 , further comprising:

if a qualifying recognized character string is not included in any character recognition data set constructed for a tier two document, identifying the tier two document for tier three processing; and

for each tier two document identified for tier three processing, performing a third tier data extraction process, the third tier data extraction process comprising:

generating a graphical representation of the tier two document at a workstation; and

for each required invoice data element for which the second tier data extraction process failed to write an output data value to the output data structure:

prompting for user input of an output data value;

receiving user input of the output data value from the workstation; and

writing, to the output data structure, the output data value received from the workstation in association with identification of the required invoice data element.

11. The method of claim 10 , wherein:

at least one positional element value includes an abscissa value and an ordinate value defining a Cartesian coordinate within the graphical representation of the output document object at which an origin of the text string is positioned; and

a qualifying text string is a text string of a print element that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

12. The method of claim 11 , wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

13. The method of claim 10 , wherein:

at least one positional element value includes:

a reference to a second invoice data element;

an abscissa value; and

an ordinate value;

wherein: i) the abscissa value added to an abscissa value of the second invoice data element; and ii) the ordinate value added to an ordinate value of the second invoice data element define a Cartesian coordinate within the graphical representation of the output document object at which an origin of the text string is positioned; and

wherein a qualifying text string is a text string of a print element that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

14. The method of claim 13 , wherein a qualifying recognized character string is a recognized character string of a character recognition data set that includes a position identifier value identifying a position within the graphical representation of the output document object that is within a predetermined displacement from the Cartesian coordinate.

15. The method of claim 10 , wherein:

the graphic image of at least one print component includes a pixelized representation of at least one character; and

the recognized character string of at least one character recognition data set includes characters matching characters of the pixelized representation of at least one character.

16. The method of claim 10 , further comprising:

crediting an account for a first charge in the event all required invoice data elements are written to the output data structure by the first tier data extraction process;

crediting the account for a second charge, different than the first charge, in the event:

any required invoice data elements are written to the output data structure by the second tier data extraction process; and

the output document object is not identified for tier three processing; and

crediting the account for a third charge, different than both the first charge and the second charge, in the event the output document object is identified for the tier three processing.

Assignments (5)
RELEASE OF SECURITY INTEREST IN REEL/FRAME: 040882/0908 Recorded May 13, 2022
From: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
To: BOTTOMLINE TECHNOLOGIES (DE), INC.
Reel/Frame 060063/0701 →
SECURITY INTEREST Recorded May 13, 2022
From: BOTTOMLINE TECHNOLOGIES, INC.
To: ARES CAPITAL CORPORATION
Reel/Frame 060064/0275 →
CHANGE OF NAME Recorded Mar 19, 2021
From: BOTTOMLINE TECHNOLOGIES (DE), INC.
To: BOTTOMLINE TECHNLOGIES, INC.
Reel/Frame 055661/0461 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Dec 12, 2016
From: BOTTOMLINE TECHNOLOGIES (DE), INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 040882/0908 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2008
From: GANGAI, DANIEL P.
To: BOTTOMLINE TECHNOLOGIES (DE) INC.
Reel/Frame 020818/0159 →
Continuity (3)
Continuation In Part 11821034 · Jun 21, 2007
Provisional Application 60923816 · Apr 16, 2007
Related Publication 20080252924A1 · Oct 16, 2008