IP Library › Granted Patent US 12,260,419
Granted Patent B2
US 12,260,419 · App. 17/181,519 · Granted Mar 25, 2025

Efficient data processing to identify information and reformat data files, and applications thereof

Inventors: Carlos Vera-Ciro (Madison, WI); Robert Raymond Lindner (Fitchburg, WI)
Assignee: VEDA Data Solutions, Inc.
G06Q30/0201G06F16/951G06F40/295G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,419
App. No.
17/181,519
Granted
Mar 25, 2025
Kind
B2
Abstract

The present disclosure is directed to systems and methods for identifying demographic information in a data file. The method may include: receiving the data file containing a plurality of fields of demographic information from a third-party, the data file having inconsistent or mislabeled nomenclatures for one or more fields of the plurality of fields or spurious demographic information; analyzing the data file using a machine learning model trained according to other data files to distinguish between each of the plurality of fields of demographic information, the machine learning model being based on a plurality of machine learning algorithms to identify different types demographic information; generating a score indicating a probability that each of the plurality of fields of demographic information was identified correctly; and generating a revised data file labeling each of the plurality of fields of demographic information based on the identified type.

Claims (56)

1. A computer-implemented method of identifying demographic information in a data file, comprising:

training a machine learning model according to labeled, sampled training sets to identify a heading based at least in part on structure and content of the data file from a data source with information describing medical providers, the machine learning model being based on a plurality of machine learning algorithms to identify different types of demographic information;

receiving, by a processor, data files containing a plurality of fields of demographic information from a plurality of third-party sources, the data files having inconsistent or mislabeled nomenclatures with one another for one or more fields of the plurality of fields of demographic information;

analyzing, by the processor, the heading identifying a plurality of strings representing a field of demographic information in the data files using the machine learning model, wherein the analyzing is performed across two or more fields of demographic information to identify a relationship between headings, wherein the identification of the relationship further comprises:

determining types of demographic information;

determining whether the headings should be grouped as pairs; and

determining whether entries of each of the headings describe a same medical provider; and

generating, by the processor, a score indicating a probability that each of the plurality of fields of demographic information was identified correctly, wherein the generating the score further comprises:

generating a baseline score for each of the plurality of fields of demographic information; and

adjusting the baseline score to increase the score when the heading and content of the field of demographic information match, or to decrease the score when the heading and the content of the field of demographic information do not match;

generating, by the processor, a revised data file labeling each of the plurality of fields of demographic information based on the identified type; and

inserting, by the processor and in the revised data file, missing fields of demographic information based on the identified type of demographic information.

2. The method of claim 1 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

3. The method of claim 1 , wherein analyzing the data file comprises analyzing a shape of each of the plurality of fields of demographic information to identify the different types of demographic information.

4. The method of claim 1 , wherein analyzing the data file comprises analyzing metadata of each of the plurality of fields of demographic information to identify the different types of demographic information.

5. The method of claim 4 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

6. The method of claim 1 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the method further comprises cross-checking at least one of the plurality of fields of demographic information against known demographic information.

7. The method of claim 1 , further comprising transmitting the revised data file to a third-party.

8. The method of claim 1 , wherein the analyzing comprises distinguishing between each of the plurality of fields of demographic information based on a position of the heading in the data files with respect to another position of another heading in the data files.

9. A system for identifying demographic information in a data file, comprising:

a memory that stores instructions for identifying the demographic information in the data file; and

a processor configured to execute the instructions that cause the processor to:

train a machine learning model according to labeled, sampled training sets to identify a heading based at least in part on structure and content of the data file from a data source with information describing medical providers, the machine learning model being based on a plurality of machine learning algorithms to identify different types of demographic information;

receive, by the processor, data files containing a plurality of fields of demographic information from a plurality of third-party sources, the data files having inconsistent or mislabeled nomenclatures with one another for one or more fields of the plurality of fields of demographic information;

analyze, by the processor, the heading identifying a plurality of strings representing a field of demographic information in the data files using the machine learning model, wherein the analyze is performed across two or more fields of demographic information to identify a relationship between headings, wherein the identification of the relationship further comprises the instructions that cause the processor to:

determine types of demographic information;

determine whether the headings should be grouped as pairs;

determine whether entries of each of the headings describe a same medical provider; and

generate, by the processor, a score indicating a probability that each of the plurality of fields of demographic information was identified correctly, wherein the generate the score further comprises the instructions that cause the processor to:

generate a baseline score for each of the plurality of fields of demographic information; and

adjust the baseline score to increase the score when the heading and content of the field of demographic information match, or to decrease the score when the heading and the content of the field of demographic information do not match;

generate, by the processor, a revised data file labeling each of the plurality of fields of demographic information based on the identified type; and

insert, by the processor and in the revised data file, missing fields of demographic information based on the identified type of demographic information.

10. The system of claim 9 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

11. The system of claim 9 , wherein analyzing the data file comprises analyzing metadata of each of the plurality of fields of demographic information to identify the different types of demographic information.

12. The system of claim 11 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

13. The system of claim 9 , wherein analyzing the data file comprises analyzing each nomenclature to identify the different types of demographic information.

14. The system of claim 9 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the instructions further cause the processor to cross-check at least one of the plurality of fields of demographic information against known demographic information.

15. The system of claim 9 , wherein the instructions further cause the processor to transmit the revised data file to a third-party.

16. A non-transitory program storage device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations, the operations comprising:

training a machine learning model according to labeled, sampled training sets to identify a heading based at least in part on structure and content of a data file from a data source with information describing medical providers, the machine learning model being based on a plurality of machine learning algorithms to identify different types of demographic information;

receiving data files containing a plurality of fields of demographic information from a plurality of third-party sources, the data files having inconsistent or mislabeled nomenclatures with one another for one or more fields of the plurality of fields of demographic information;

analyzing the heading identifying a plurality of strings representing a field of demographic information in the data files using the machine learning model, wherein the analyzing is performed across two or more fields of demographic information to identify a relationship between headings, wherein the identification of the relationship further comprises:

determining types of demographic information;

determining whether the headings should be grouped as pairs; and

determining whether entries of each of the headings describe a same medical provider; and

generating a score indicating a probability that each of the plurality of fields of demographic information was identified correctly, wherein the generating the score further comprises:

generating a baseline score for each of the plurality of fields of demographic information; and

adjusting the baseline score to increase the score when the heading and content of the field of demographic information match, or to decrease the score when the heading and the content of the field of demographic information do not match;

generating a revised data file labeling each of the plurality of fields of demographic information based on the identified type; and

inserting in the revised data file, missing fields of demographic information based on the identified type of demographic information.

17. The non-transitory program storage device of claim 16 , wherein analyzing the data file comprises analyzing semantic content of each of the plurality of fields of demographic information to identify the different types of demographic information.

18. The non-transitory program storage device of claim 16 , wherein analyzing the data file comprises analyzing a shape of each of the plurality of fields of demographic information to identify the different types of demographic information.

19. The non-transitory program storage device of claim 16 , wherein analyzing the data file comprises analyzing metadata of each of the plurality of fields of demographic information to identify the different types of demographic information.

20. The non-transitory program storage device of claim 19 , wherein the metadata includes each nomenclature of each of the plurality of fields of demographic information.

21. The non-transitory program storage device of claim 16 , wherein, in response to identifying different ones of the plurality of fields of demographic information, the operations further comprise cross-checking at least one of the plurality of fields of demographic information against known demographic information.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: VERA-CIRO, CARLOS; LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 055363/0094 →
Continuity (2)
Continuation 16668565 · Oct 30, 2019
Related Publication 20210174380A1 · Jun 10, 2021
References Cited (43)
US 8780195B1 · Rubin · 2014 [cited by applicant]
US 9529359B1 · Annan et al. · 2016 [cited by applicant]
US 9626432B2 · Acevedo Arizpe · 2017 [cited by examiner]
US 10366351B2 · Whittier · 2019 [cited by examiner]
US 10387825B1 · Canavor et al. · 2019 [cited by applicant]
US 10963378B2 · Somech · 2021 [cited by examiner]
US 11120899B1 · Rai · 2021 [cited by examiner]
US 11862305B1 · Sethi · 2024 [cited by examiner]
US 20030182310A1 · Charnock et al. · 2003 [cited by applicant]
US 20070282681A1 · Shubert et al. · 2007 [cited by applicant]
US 20090119576A1 · Pepper et al. · 2009 [cited by applicant]
US 20120053959A1 · Sengupta · 2012 [cited by examiner]
US 20120209795A1 · Glickman · 2012 [cited by applicant]
US 20120226719A1 · Sewall · 2012 [cited by applicant]
US 20120272259A1 · Cortes et al. · 2012 [cited by applicant]
US 20140195544A1 · Whitman · 2014 [cited by applicant]
US 20140249865A1 · Ghani · 2014 [cited by examiner]
US 20160019197A1 · Lasi · 2016 [cited by examiner]
US 20160135706A1 · Sullivan et al. · 2016 [cited by applicant]
US 20160147943A1 · Ash · 2016 [cited by examiner]
US 20160173827A1 · Dannan et al. · 2016 [cited by applicant]
US 20160236790A1 · Knapp et al. · 2016 [cited by applicant]
US 20160283350A1 · Akbulut · 2016 [cited by examiner]
US 20160288905A1 · Gong et al. · 2016 [cited by applicant]
US 20160314123A1 · Ramachandran · 2016 [cited by examiner]
US 20170017638A1 · Satyavarta et al. · 2017 [cited by applicant]
US 20170161725A1 · Hosp et al. · 2017 [cited by applicant]
US 20170286622A1 · Cox et al. · 2017 [cited by applicant]
US 20180060418A1 · Robichaud · 2018 [cited by applicant]
US 20180121514A1 · Reisz · 2018 [cited by examiner]
US 20190113905A1 · Corr · 2019 [cited by examiner]
US 20190213408A1 · Cali · 2019 [cited by examiner]
US 20190287685A1 · Wu · 2019 [cited by examiner]
US 20190311299A1 · Lindner · 2019 [cited by applicant]
US 20190392075A1 · Han et al. · 2019 [cited by applicant]
US 20200004765A1 · Sørensen · 2020 [cited by examiner]
US 20200273570A1 · Subramanian et al. · 2020 [cited by applicant]
US 20210174380A1 · Vera-Ciro et al. · 2021 [cited by applicant]
CN 112817847A · 2021 [cited by applicant]
WO WO2007149216A2 · 2007 [cited by applicant]
Mallieswari et al., Effect of Machine Learning in Healthcare Industry with reference to Artificial Intelligence, International Journal of Advanced Research in Computer Engineering & Technology, vol. 8, pp. 10-13 (Year: … [cited by examiner]
Huddar et al., Predicting Complications in Critical Care Using Heterogeneous Clinical Data, 2016, IEEE Access, vol. 4, 7988-8001 (Year: 2016). [cited by examiner]
Meystre et a. BMC Medical Research Methodology (Year: 2010). [cited by examiner]