IP Library Granted Patent US 9,020,988
Granted Patent B2
US 9,020,988 · App. 13/732,145 · Granted Apr 28, 2015

Database aggregation of purchase data

Inventors: Jeffrey David Rubenstein (Boca Raton, FL); Wagner Figueiredo Dosanjos (Clearwater, FL)
Assignee: SmartProcure, Inc.
G06F17/30489G06Q30/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,020,988
App. No.
13/732,145
Granted
Apr 28, 2015
Kind
B2
Abstract

A big data database managed by a procurement service aggregates purchase data received from federal, state and local government agencies through Freedom of Information Act requests, state public records requests and private sector business entities. An automated system processes a vast amount of purchase data files acquired from numerous different agencies through a number of different transports, on a variety of different media, included within several different file formats. A best match is selected for each acquired file with one of a multitude of configuration files available to process the purchase files. The file is then processed with the selected configuration file and aggregated into the database. The database is then made available to customers of the procurement service for search, reports and analysis purposes.

Claims (124)

1. An automated method of aggregating data for a database comprising:

acquiring a purchase data file from one of a multiplicity purchasing agencies;

converting the purchase data file into a multiplicity of text files using a multiplicity of configuration files, each configuration file having

a strategy for converting the purchase data file into an intermediate text file having lines of text,

instruction lines for processing the lines of text of the intermediate text file into an intermediate data file having a plurality of fields, and

field definitions for validating at least some of the plurality of fields of the intermediate data file;

for each of the multiplicity of configuration files, processing the intermediate text file into the intermediate data file having the plurality of fields and analyzing the processing to produce processing results;

selecting a best match configuration file of the multiplicity of configuration files determined to have a higher level of processing results;

testing the best match configuration file for a minimum level of acceptable data;

parsing the intermediate text file of the best match configuration file into an intermediate annotated data file using the best match configuration file;

normalizing the intermediate annotated data file to produce a normalized intermediate annotated data file;

loading the normalized intermediate annotated data file into a backend database; and

publishing changes in the backend database to a customer facing database, wherein the minimum level of acceptable data includes a population of at least one of the following fields: purchase order number, vendor name, date, quantity, description and amount.

2. The method according to claim 1 wherein the purchase data file is acquired in a file format including one of csv, docx, excel, html, pdf, pipe separated, rtf, tab separated, text, text fixed and jpeg.

3. The method according to claim 1 wherein each configuration file includes a plurality of tag definitions indicative of a third party application used to generate purchase data files and

the analyzing the processing to produce processing results determines a number of the plurality of tag definitions occurring in the intermediate data file and the process results include the number of the plurality of tag definitions occurring in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of the plurality of tag definitions occurring in the intermediate data file.

4. The method according to claim 1 wherein

the analyzing the processing to produce processing results determines a number of fields populated in the intermediate data file and the process results include the number of fields populated in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of fields populated in the intermediate data file.

5. The method according to claim 1 wherein

the processing the intermediate text file into the intermediate data file further includes validating the fields using the field definitions,

the analyzing the processing to produce processing results determines a number of validated fields populated in the intermediate data file and the process results include the number of validated fields populated in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of validated fields populated in the intermediate data file.

6. The method according to claim 1 wherein

the analyzing the processing to produce processing results determines a number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file and the process results include the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file.

7. The method according to claim 1 wherein

the analyzing the processing to produce processing results determines a number of lines of instructions in a configuration file utilized by the processing of the intermediate text file into the intermediate data file and the process results include the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file.

8. The method according to claim 1 wherein

each configuration file includes a plurality of tag definitions indicative of a third party application used to generate purchase data files,

the processing the intermediate text file into the intermediate data file further includes validating the fields using the field definitions,

the analyzing the processing to produce processing results determines

a number of the plurality of tag definitions occurring in the intermediate data file,

a number of fields populated in the intermediate data file,

a number of validated fields populated in the intermediate data file,

a number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

a number of lines of instructions in a configuration file utilized by the processing of the intermediate text file into the intermediate data file,

the process result includes

the number of the plurality of tag definitions occurring in the intermediate data file,

the number of fields populated in the intermediate data file,

the number of validated fields populated in the intermediate data file,

the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to at least two of

the number of the plurality of tag definitions occurring in the intermediate data file,

the number of fields populated in the intermediate data file,

the number of validated fields populated in the intermediate data file,

the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file.

9. The method according to claim 1 wherein the intermediate annotated data file includes vendor name data represented in one of a of a multitude of textual representations and the normalizing includes:

comparing the vendor name with a vendor database having a multiplicity of formats associated with a multiplicity of vendor names;

selecting a best match vendor name from the comparing using a heuristic algorithm; and

inserting the best match vendor name into the annotated data file.

10. The method according to claim 1 wherein the testing the best match configuration file includes determining if the intermediate data file includes fields for purchase order number, vendor name and date.

11. An apparatus for automated aggregation of data for a database comprising:

an acquisition module for acquiring a purchase data file from one of a multiplicity purchasing agencies;

a configuration file selection means comprised within a computer system, the configuration file selection means for converting the purchase data file into a multiplicity of text files using a multiplicity of configuration files, each configuration file having

a strategy for converting the purchase data file into an intermediate text file having lines of text,

instruction lines for processing the lines of text of the intermediate text file into an intermediate data file having a plurality of fields, and

field definitions for validating at least some of the plurality of fields of the intermediate data file;

for each of the multiplicity of configuration files, processing the intermediate text file into the intermediate data file having the plurality of fields and analyzing the processing to produce processing results;

the configuration file selection means further for selecting a best match configuration file of the multiplicity of configuration files determined to have a higher level of processing results;

a parsing module for parsing the intermediate text file of the best match configuration file into an intermediate annotated data file using the best match configuration file;

a load data module for loading the intermediate annotated data file into a backend database;

a normalization module for normalizing the intermediate annotated data file to produce a normalized intermediate annotated data file; and

a publish module for publishing changes in the backend database to a customer facing database.

12. The apparatus according to claim 11 wherein the purchase data file is acquired in a file format including one of csv, docx, excel, html, pdf, pipe separated, rtf, tab separated, text, text fixed and jpeg.

13. The apparatus according to claim 11 wherein

the configuration file includes a plurality of tag definitions indicative of a third party application used to generate purchase data files,

the processing the intermediate text file into the intermediate data file further includes validating the fields using the field definitions,

the analyzing the processing to produce processing results determines

a number of the plurality of tag definitions occurring in the intermediate data file,

a number of fields populated in the intermediate data file,

a number of validated fields populated in the intermediate data file,

a number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

a number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file,

the process results include

the number of the plurality of tag definitions occurring in the intermediate data file,

the number of fields populated in the intermediate data file,

the number of validated fields populated in the intermediate data file,

the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to at least two of

the number of the plurality of tag definitions occurring in the intermediate data file,

the number of fields populated in the intermediate data file,

the number of validated fields populated in the intermediate data file,

the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file.

14. The apparatus according to claim 11 wherein the intermediate annotated data file includes vendor name data represented in one of a multitude of textual representations and the normalization module is further for:

comparing the vendor name with a vendor database having a multiplicity of formats associated with a multiplicity of vendor names;

selecting a best match vendor name from the comparing using a heuristic algorithm; and

inserting the best match vendor name into the annotated data file.

15. A durable, non-transitory computer readable storage medium comprising a computer program which instructs a computer to perform an automated method of aggregating data for a database, the method comprising:

acquiring a purchase data file from one of a multiplicity purchasing agencies;

converting the purchase data file into a multiplicity of text files using a multiplicity of configuration files, each configuration file having

a strategy for converting the purchase data file into an intermediate text file having lines of text,

instruction lines for processing the lines of text of the intermediate text file into an intermediate data file having a plurality of fields, and

field definitions for validating at least some of the plurality of fields of the intermediate data file;

for each of the multiplicity of configuration files, processing at least a portion of the intermediate text file into the intermediate data file having the plurality of fields and analyzing the processing to produce processing results;

selecting a best match configuration file of the multiplicity of configuration files determined to have a higher level of processing results;

testing the best match configuration file for a minimum level of acceptable data;

parsing the intermediate text file of the best match configuration file into an intermediate annotated data file using the best match configuration file;

loading the intermediate annotated data file into a backend database;

normalizing the intermediate annotated data file to produce a normalized intermediate annotated data file; and

publishing changes in the backend database to a customer facing database, wherein the minimum level of acceptable data includes a population of at least one of the following fields: purchase order number, vendor name, date, quantity, description and amount.

16. The computer program according to claim 15 wherein the purchase data file is acquired in a file format including one of csv, docx, excel, html, pdf, pipe separated, rtf, tab separated, text, text fixed and jpeg.

17. The computer program according to claim 15 wherein a configuration file includes a plurality of tag definitions indicative of a third party application used to generate purchase data files and

the analyzing the processing to produce processing results determines a number of the plurality of tag definitions occurring in the intermediate data file and the process results include the number of the plurality of tag definitions occurring in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of the plurality of tag definitions occurring in the intermediate data file.

18. The computer program according to claim 15 wherein

the analyzing the processing to produce processing results determines a number of fields populated in the intermediate data file and the process results include the number of fields populated in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of fields populated in the intermediate data file.

19. The computer program according to claim 15 wherein

the processing of the at least the portion of the intermediate text file into the intermediate data file further includes validating the fields using the field definitions,

the analyzing the processing to produce processing results determines a number of validated fields populated in the intermediate data file and the process results include the number of validated fields populated in the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of validated fields populated in the intermediate data file.

20. The computer program according to claim 15 wherein

the analyzing the processing to produce processing results determines a number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file and the process results include the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of lines of text in the intermediate text file utilized by the processing of the intermediate text file into the intermediate data file.

21. The computer program according to claim 15 wherein

the analyzing the processing to produce processing results determines a number of lines of instructions in a configuration file utilized by the processing of the intermediate text file into the intermediate data file and the process results include the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file, and

the selecting the best match configuration file of the multiplicity of configuration files selects the best match configuration file in response to the number of lines of instructions in the configuration file utilized by the processing of the intermediate text file into the intermediate data file.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Jun 28, 2024
From: FIFTH THIRD BANK, NATIONAL ASSOCIATION
To: SMARTPROCURE, INC.
Reel/Frame 067872/0693 →
SECURITY INTEREST Recorded Jul 1, 2021
From: SMARTPROCURE, INC
To: FIFTH THIRD BANK
Reel/Frame 056735/0744 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: SMARTPROCURE, LLC
To: SMARTPROCURE, INC.
Reel/Frame 034477/0039 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2013
From: RUBENSTEIN, JEFFREY DAVID; DOSANJOS, WAGNER FIGUEIREDO
To: SMARTPROCURE, LLC
Reel/Frame 029773/0664 →
Continuity (1)
Related Publication 20140188948A1 · Jul 3, 2014