IP Library › Granted Patent US 12,443,614
Granted Patent B2
US 12,443,614 · App. 18/114,075 · Granted Oct 14, 2025

Efficient column detection using sequencing, and applications thereof

Inventors: Robert Raymond Lindner (Fitchburg, WI); Carlos Vera-Ciro (Madison, WI)
Assignee: VEDA Data Solutions, Inc.
G06F16/248G06F16/2237
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,614
App. No.
18/114,075
Granted
Oct 14, 2025
Kind
B2
Abstract

The present disclosure is directed to systems and methods for identifying demographic information in a data file. The method uses evidence for a label within the column itself. In addition, the method uses likelihoods that a first label may exist at a particular location in a data file with respect to a second label. Finally, the method uses likelihoods that a first label exists in the data file at a first frequency given that a second label exists in a data file in a second frequency. All based on these likelihoods, an overall likelihood of that the label configuration is correct is determined. Using that likelihood score, a nonlinear optimization algorithm is applied to identify a best fit between a group of labels and a group of columns in a data file.

Claims (94)

1. A method for determining which labels correspond to columns of a data file comprising columns and rows, the method comprising:

receiving the data file containing a plurality of fields organized as a table with a plurality of columns and a plurality of rows, the data file having inconsistent labeling for the columns;

training a model to detect a label based on data within the columns, wherein the model is trained using:

a number of Monte Carlo training sets having sample data files, and

rules for common types of demographic information;

for respective columns in the plurality of columns, selecting, from a plurality of consistent labels, a label corresponding to the respective columns of the data file using the model to detect the label based on data within the column;

for respective first and second columns in the plurality of columns, determining a column score indicating a likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determining, based on the column score, a placement score indicating a likelihood that labels from the plurality of consistent labels corresponding to the respective columns are correct, wherein for all of the labels the determining the placement score further comprises,

determining a first frequency of a first label amongst the labels,

determining a second frequency of a second label amongst the labels,

determining a frequency score indicating a likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determining the placement score based on the frequency score;

adjusting the labels corresponding to each of the respective first and second columns based on the placement score;

repeating, until the placement score converges, steps of:

for the respective first and second columns in the plurality of columns, determining the column score indicating the likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determining, based on the column score, the placement score indicating the likelihood that the labels corresponding to the respective columns are correct, wherein for all of the labels, the determining the placement score further comprises:

determining the first frequency of the first label amongst the labels,

determining the second frequency of the second label amongst the labels,

determining the frequency score indicating the likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determining the placement score based on the frequency score; and

adjusting the labels correspond to each of the respective first and second columns based on the placement score; and

generating, by a column detector on a computing device, a reformatted data file based on the adjusted labels.

2. The method of claim 1 , wherein the determining the frequency score comprises looking up a number of times the frequency score has shown up in a historical matrix.

3. The method of claim 1 , further comprising repeating steps of:

determining the first frequency of the first label among the labels,

determining the second frequency of the second label among the labels, and

determining the frequency score indicating the likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency,

wherein the steps are repeated for respective columns in the plurality of columns to determine a plurality of frequency scores, and wherein the determining the placement score further comprises determining the placement score based on the plurality of frequency scores.

4. The method of claim 1 , further comprising for respective columns in the plurality of columns, determining an evidence score indicating a likelihood that the label is correct, wherein the determining the placement score further comprises determining the placement score based on the evidence scores.

5. The method of claim 4 , wherein the data file lists health care providers.

6. The method of claim 1 , wherein the determining the placement score comprises looking up a number of times the placement score has shown up in a historical matrix.

7. The method of claim 1 , wherein the data file stores demographic information.

8. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations for determining which labels correspond to columns of a data file comprising columns and rows, the operations comprising:

receiving the data file containing a plurality of fields organized as a table with a plurality of columns and a plurality of rows, the data file having inconsistent labeling for the columns;

training a model to detect a label based on data within the columns, wherein the model is trained using:

a number of Monte Carlo training sets having sample data files, and

rules for common types of demographic information;

for respective columns in the plurality of columns, selecting, from a plurality of consistent labels, a label corresponding to the respective columns of the data file using the model to detect the label based on data within the column;

for respective first and second columns in the plurality of columns, determining a column score indicating a likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determining, based on the column score, a placement score indicating a likelihood that labels from the plurality of consistent labels corresponding to the respective columns are correct, wherein for all of the labels, the operations further comprise:

determining a first frequency of a first label amongst the labels,

determining a second frequency of a second label amongst the labels,

determining a frequency score indicating a likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determining the placement score based on the frequency score;

adjusting the labels corresponding to each of the respective columns based on the placement score;

repeating, until the placement score converges, steps of:

for the respective first and second columns in the plurality of columns, determining the column score indicating the likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determining, based on the column score, the placement score indicating the likelihood that the labels corresponding to the respective columns are correct, wherein for all of the labels, the operations further comprise:

determining the first frequency of the first label amongst the labels,

determining the second frequency of the second label amongst the labels,

determining the frequency score indicating the likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determining the placement score based on the frequency score; and

adjusting the labels correspond to each of the respective first and second columns based on the placement score; and

generating, by a column detector on the non-transitory computer-readable device, a reformatted data file based on the adjusted labels.

9. The non-transitory computer-readable device of claim 8 , wherein the determining the frequency score comprises looking up a number of times the frequency score has shown up in a historical matrix.

10. The non-transitory computer-readable device of claim 8 , wherein the operations further comprise repeating steps of:

determining the first frequency of the first label among the labels,

determining the second frequency of the second label among the labels, and

determining the frequency score indicating the likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency,

wherein the steps are repeated for respective columns in the plurality of columns to determine a plurality of frequency scores, and wherein the determining the placement score further comprises determining the placement score based on the plurality of frequency scores.

11. The non-transitory computer-readable device of claim 8 , wherein the operations further comprise for respective columns in the plurality of columns, determining an evidence score indicating a likelihood that the label is correct, wherein the determining the placement score further comprises determining the placement score based on the evidence scores.

12. The non-transitory computer-readable device of claim 8 , wherein the determining the placement score comprises looking up the number of times the placement score has shown up in a historical matrix.

13. The non-transitory computer-readable device of claim 8 , wherein the data file stores demographic information.

14. The non-transitory computer-readable device of claim 13 , wherein the data file lists health care providers.

15. The non-transitory computer-readable device of claim 8 , wherein the operations further comprise repeating, until an optima is found, the steps of:

for the respective first and second columns in the plurality of columns, determining the column score indicating the likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column,

determining, based on the column score, the placement score indicating the likelihood that the labels corresponding to the respective columns are correct, and

adjusting the labels corresponding to each of the respective first and second columns.

16. The non-transitory computer-readable device of claim 8 , wherein the adjusting which of the labels correspond to each of the respective columns to obtain the reformatted data file comprises applying a nonlinear optimization technique.

17. The non-transitory computer-readable device of claim 16 , wherein the nonlinear optimization technique includes at least one of hill climbing, stochastic hill climbing, simulated, local beams search, or genetic algorithms.

18. A system for determining which labels correspond to columns of a data file comprising columns and rows, comprising:

a processor; and

a memory with instructions stored thereon that when executed by the processor, cause the processor to:

receive the data file containing a plurality of fields organized as a table with a plurality of columns and a plurality of rows, the data file having inconsistent labeling for the columns;

train a model to detect a label based on data within the columns, wherein the model is trained using:

a number of Monte Carlo training sets having sample data files, and

rules for common types of demographic information;

for respective columns in the plurality of columns, select, from a plurality of consistent labels, a label corresponding to the respective columns of the data file using the model to detect the label based on data within the column;

for respective first and second columns in the plurality of columns, determine a column score indicating a likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determine, based on the column score, a placement score indicating a likelihood that labels from the plurality of consistent labels corresponding to the respective columns are correct, wherein for all of the labels, the instructions further cause the processor to:

determine a first frequency of a first label amongst the labels,

determine a second frequency of a second label amongst the labels,

determine a frequency score indicating a likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determine the placement score based on the frequency score;

adjust the labels corresponding to each of the respective columns based on the placement score;

repeat, until the placement score converges, steps of:

for the respective first and second columns in the plurality of columns, determine the column score indicating the likelihood that the first column has the label corresponding to the first column given that the second column has the label corresponding to the second column;

determine, based on the column score, the placement score indicating the likelihood that the labels corresponding to the respective columns are correct, wherein for all of the labels, the instructions further cause the processor to:

determine the first frequency of the first label amongst the labels,

determine the second frequency of the second label amongst the labels,

determine the frequency score indicating the likelihood that the first label occurs at the first frequency given that the second label occurs at the second frequency, and

determine the placement score based on the frequency score; and

adjust the labels correspond to each of the respective first and second columns based on the placement score; and

generate, by a column detector on the system, a reformatted data file based on the adjusted labels.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
Continuity (2)
Provisional Application 63268539 · Feb 25, 2022
Related Publication 20230273934A1 · Aug 31, 2023
References Cited (11)
US 20120166927A1 · Shearer · 2012 [cited by examiner]
US 20130124960A1 · Velingkar et al. · 2013 [cited by applicant]
US 20160103863A1 · Becker · 2016 [cited by examiner]
US 20170060919A1 · Ramachandran · 2017 [cited by examiner]
US 20180060418A1 · Robichaud · 2018 [cited by applicant]
US 20190311299A1 · Lindner · 2019 [cited by applicant]
US 20190392075A1 · Han et al. · 2019 [cited by applicant]
US 20200321113A1 · Neumann · 2020 [cited by examiner]
US 20210174380A1 · Vera-Ciro et al. · 2021 [cited by applicant]
US 20220309390A1 · Arnautov · 2022 [cited by examiner]
International Search Report and Written Opinion directed to related International Application No. PCT /US2023/063200, mailed Jun. 7, 2023; 6 pages. [cited by applicant]