IP Library › Patent Application 16668565
Patent Application
App. No. 16/668,565

EFFICIENT DATA PROCESSING TO IDENTIFY INFORMATION AND REFORMANT DATA FILES, AND APPLICATIONS THEREOF

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/668,565
Abstract

The present disclosure is directed to systems and methods for identifying demographic information in a data file. The method may include: receiving the data file containing a plurality of fields of demographic information from a third-party, the data file having inconsistent or mislabeled nomenclatures for one or more fields of the plurality of fields or spurious demographic information; analyzing the data file using a machine learning model trained according to other data files to distinguish between each of the plurality of fields of demographic information, the machine learning model being based on a plurality of machine learning algorithms to identify different types demographic information; generating a score indicating a probability that each of the plurality of fields of demographic information was identified correctly; and generating a revised data file labeling each of the plurality of fields of demographic information based on the identified type.

Claims (34)

1 . A computer-implemented method of identifying demographic information in a data file, comprising:

receiving the data file containing a plurality of fields of demographic information from a third-party, the data file having inconsistent or mislabeled nomenclatures for one or more fields of the plurality of fields or spurious demographic information;

analyzing the data file using a machine learning model trained according to other data files to distinguish between each of the plurality of fields of demographic information, the machine learning model being based on a plurality of machine learning algorithms to identify different types demographic information;

generating a score indicating a probability that each of the plurality of fields of demographic information was identified correctly; and

generating a revised data file labeling each of the plurality of fields of demographic information based on the identified type.

2 . The method of claim 1 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

3 . The method of claim 1 , wherein analyzing the data file comprises analyzing a shape of each of the plurality of fields of demographic information to identify the different types of demographic information.

4 . The method of claim 1 , wherein analyzing the data file comprises analyzing metadata of each of the plurality of fields of demographic information to identify the different types of demographic information.

5 . The method of claim 4 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

6 . The method of claim 1 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the method further comprises cross-checking at least one of the plurality of fields of demographic information against known demographic information.

7 . The method of claim 1 , further comprising transmitting the revised data file to the third-party.

8 . A system for identifying demographic information in a data file, comprising:

a memory that stores instructions for identifying the demographic information in the data file; and

a processor configured to execute the instructions that cause the processor to:

receive the data file containing a plurality of fields of demographic information from a third-party, the data file having inconsistent or mislabeled nomenclatures for one or more fields of the plurality of fields or spurious demographic information;

analyze the data file using a machine learning model trained according to other data files to distinguish between each of the plurality of fields of demographic information, the machine learning model being based on a plurality of machine learning algorithms to identify different types demographic information;

generate a score indicating a probability that each of the plurality of fields of demographic information was identified correctly; and

generate a revised data file labeling each of the plurality of fields of demographic information based on the identified type.

9 . The system of claim 8 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

10 . The system of claim 8 , wherein analyzing the data file comprises analyzing a shape of each of the plurality of fields of demographic information to identify the different types of demographic information.

11 . The system of claim 10 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

12 . The system of claim 8 , wherein analyzing the data file comprises analyzing each nomenclature to identify the different types of demographic information.

13 . The system of claim 8 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the instructions further cause the processor to cross-check at least one of the plurality of fields of demographic information against known demographic information.

14 . The system of claim 8 , wherein the instructions further cause the processor to transmit the revised data file to the third-party.

15 . non-transitory program storage device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform a method, the method comprising:

receiving the data file containing a plurality of fields of demographic information from a third-party, the data file having inconsistent or mislabeled nomenclatures for one or more fields of the plurality of fields or spurious demographic information;

analyzing the data file using a machine learning model trained according to other data files to distinguish between each of the plurality of fields of demographic information, the machine learning model being based on a plurality of machine learning algorithms to identify different types demographic information;

generating a score indicating a probability that each of the plurality of fields of demographic information was identified correctly; and

generating a revised data file labeling each of the plurality of fields of demographic information based on the identified type.

16 . The method of claim 15 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

17 . The method of claim 15 , wherein analyzing the data file comprises analyzing a shape of each of the plurality of fields of demographic information to identify the different types of demographic information.

18 . The method of claim 15 , wherein analyzing the data file comprises analyzing metadata of each of the plurality of fields of demographic information to identify the different types of demographic information.

19 . The method of claim 18 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

20 . The method of claim 15 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the method further comprises cross-checking at least one of the plurality of fields of demographic information against known demographic information.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: VERA-CIRO, CARLOS; LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 050883/0151 →