IP Library › Granted Patent US 10,929,406
Granted Patent B2
US 10,929,406 · App. 15/335,831 · Granted Feb 23, 2021

Systems and methods for a self-services data file configuration with various data sources

Inventors: Jane Cook (Mountain Lakes, NJ); Sachin Jadhav (Phoenix, AZ); Yogaraj Jayaprakasam (Phoenix, AZ); Deepak Narayanan (Phoenix, AZ); Rahul Shaurya (Phoenix, AZ)
Assignee: American Express Travel Related Services Company, Inc
G06F16/2457G06F16/182G06F16/211G06F16/221G06F16/2282G06F16/2471
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,929,406
App. No.
15/335,831
Filed
Oct 27, 2016
Granted
Feb 23, 2021
Kind
B2
Art Unit
2163
USPC
707/809
Abstract

A system for generating and delivering custom data sets in a big data environment may receive a preselected schema that identifies a plurality of columns from a plurality of data sources for inclusion in an output data file. The system reads data from the data sources to generate a data file containing a big data table. The system monitors the plurality of data sources to detect that the data sources have been ingested into a data storage system. The data file is read and a column is filtered from the data file to generate the output data file in response to the preselected schema excluding the column. The output data file is transferred to a client device.

Claims (38)

1. A method of generating and delivering custom data sets in a data environment, comprising:

receiving a preselected schema that identifies a plurality of columns from a plurality of data sources for inclusion in a customized output data file;

reading data from the plurality of data sources to generate a data file containing a data table by extracting data from a data storage system, staging the data in a staging table, preprocessing the staging table to generate the data file, and storing the data file in a use case folder;

monitoring the plurality of data sources to detect that the plurality of data sources have been ingested into the data storage system by listening to a messaging queue to detect a timestamp of the data file, comparing the timestamp of the data file to a timestamp of a trigger in a trigger table to determine that a data source associated with the trigger has been ingested into the data storage system, and completing a data readiness check to determine that individual ones of the plurality of data sources have been ingested more recently than the timestamp of the data file;

reading the data file and filtering a column from the data file in response to the data readiness check and the preselected schema excluding the column to generate the customized output data file by executing a first query against the data file to generate a first schema and registering the first schema in a temporary table, wherein the customized output data file comprises the plurality of columns from the plurality of data sources; and

transferring the customized output data file to a client device.

2. The method of claim 1 , wherein the transferring the output data file to the client device comprises transferring the data file over a secure file transfer channel.

3. The method of claim 1 , further comprising creating a formatted and ordered string containing content in response to the reading the data file.

4. The method of claim 3 , further comprising writing the customized output data file to a distributed file system.

5. A computer-based system, comprising:

a processor;

a tangible, non-transitory memory configured to communicate with the processor, the tangible, non-transitory memory having instructions stored thereon that, in response to execution by the processor, cause the computer-based system to perform operations comprising:

receiving a preselected schema that identifies a plurality of columns from a plurality of data sources for inclusion in a customized output data file;

reading data from the plurality of data sources to generate a data file containing a data table by extracting data from a data storage system, staging the data in a staging table, preprocessing the staging table to generate the data file, and storing the data file in a use case folder;

monitoring the plurality of data sources to detect that the plurality of data sources have been ingested into the data storage system by listening to a messaging queue to detect a timestamp of the data file, comparing the timestamp of the data file to a timestamp of a trigger in a trigger table to determine that a data source associated with the trigger has been ingested into the data storage system, and completing a data readiness check to determine that individual ones of the plurality of data sources have been ingested more recently than the timestamp of the data file;

reading the data file and filtering a column from the data file in response to the data readiness check and the preselected schema excluding the column to generate the customized output data file by executing a first query against the data file to generate a first schema and registering the first schema in a temporary table, wherein the customized output data file comprises the plurality of columns from the plurality of data sources; and

transferring the customized output data file to a client device.

6. The computer-based system of claim 5 , wherein the transferring the customized output data file to the client device further comprises transferring the data file over a secure file transfer channel.

7. The computer-based system of claim 5 , wherein the reading the data file further comprises:

executing a first query against the data file to generate a first schema; and

registering the first schema in a temporary table.

8. The computer-based system of claim 5 , further comprising creating a formatted and ordered string containing content in response to the reading the data file.

9. The computer-based system of claim 5 , further comprising writing the customized output data file to a distributed file system.

10. An article of manufacture including a non-transitory, tangible computer readable storage medium having instructions stored thereon that, in response to execution by a computer-based system, cause the computer-based system to perform operations comprising:

receiving, by the computer-based system, a preselected schema that identifies a plurality of columns from a plurality of data sources for inclusion in a customized output data file;

reading, by the computer-based system, data from the plurality of data sources to generate a data file containing a data table by extracting data from a data storage system, staging the data in a staging table, preprocessing the staging table to generate the data file, and storing the data file in a use case folder;

monitoring, by the computer-based system, the plurality of data sources to detect that the plurality of data sources have been ingested into the data storage system by listening to a messaging queue to detect a timestamp of the data file, comparing the timestamp of the data file to a timestamp of a trigger in a trigger table to determine that a data source associated with the trigger has been ingested into the data storage system, and completing a data readiness check to determine that individual ones of the plurality of data sources have been ingested more recently than the timestamp of the data file;

reading the data file and filtering a column from the data file in response to the data readiness check and the preselected schema excluding the column to generate the customized output data file by executing a first query against the data file to generate a first schema and registering the first schema in a temporary table, wherein the customized output data file comprises the plurality of columns from the plurality of data sources; and

transferring, by the computer-based system, the customized output data file to a client device.

11. The method of claim 1 , wherein extracting data from the data storage system further comprises extracting the data at predetermined times by creating a schedule of Cron jobs.

12. The computer-based system of claim 5 , wherein extracting data from the data storage system further comprises extracting the data at predetermined times by creating a schedule of Cron jobs.

13. The article of claim 10 , wherein extracting data from the data storage system further comprises extracting the data at predetermined times by creating a schedule of Cron jobs.

14. The computer-based system of claim 5 , further comprising writing the customized output data file to a disk associated with the data storage system.

15. The method of claim 1 , further comprising writing the customized output data file to a disk associated with the data storage system.

16. The article of claim 10 , further comprising writing the customized output data file to a disk associated with the data storage system.

17. The article of claim 10 , wherein the transferring the output data file to the client device comprises transferring the data file over a secure file transfer channel.

18. The article of claim 10 , further comprising creating a formatted and ordered string containing content in response to the reading the data file.

19. The article of claim 18 , further comprising writing the customized output data file to a distributed file system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2016
From: COOK, JANE; JADHAV, SACHIN; NARAYANAN, DEEPAK; SHAURYA, RAHUL; JAYAPRAKASAM, YOGARAJ
To: AMERICAN EXPRESS TRAVEL RELATED SERVICES COMPANY, INC.
Reel/Frame 040150/0486 →
Continuity (1)
Related Publication 20180121519A1 · May 3, 2018