IP Library › Granted Patent US 12,572,853
Granted Patent B2
US 12,572,853 · App. 18/094,120 · Granted Mar 10, 2026

Training machine learning algorithms with temporally variant personal data, and applications thereof

Inventor: Robert Raymond Lindner (Fitchburg, WI)
Assignee: H1 Insights, Inc.
G06N20/00G06F16/2365G06F16/248G06F21/6209G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,853
App. No.
18/094,120
Filed
Jan 6, 2023
Granted
Mar 10, 2026
Kind
B2
Art Unit
2159
USPC
706/12
Abstract

To train models, training data is needed. As personal data changes over time, the training data can get stale, obviating its usefulness in training the model. Embodiments deal with this by developing a database with a running log specifying how each person's data changes at the time. When data is ingested, it may not be normalized. To deal with this, embodiments clean the data to ensure the ingested data fields are normalized. Finally, the various tasks needed to train the model and solve for accuracy of personal data can quickly become cumbersome to a computing device. They can conflict with one another and compete inefficiently for computing resources, such as processor power and memory capacity. To deal with these issues, a scheduler is employed to queue the various tasks involved.

Claims (40)

1 . A computer-implemented method for linking ingested data, the method comprising:

accessing, by one or more computing devices, data records of individuals;

parsing, by the one or more computing devices, the data records to locate each individual's demographic information;

assigning, by the one or more computing devices, the demographic information into predetermined categories;

comparing, by the one or more computing devices, each record against first categorized records using a pair-wise function to determine if they are the same;

training, by the one or more computing devices, a neural network using a training set, wherein the training set is retrieved from an database at a particular time an accuracy of the database was verified;

calculating, using the trained neural network and by the one or more computing devices, a similarity score for each pair of records, wherein the similarity score is a ratio based on how many categories match for each data pair of records;

determining, by the one or more computing devices, whether the similarity score meets or exceeds a similarity score threshold;

based on determining the similarity score meets or exceeds the similarity score threshold, linking, by the one or more computing devices, the data pair in a group;

determining, by the one or more computing devices, a most prevalent identity within the group, comprising:

comparing each record against second categorized records to determine the most prevalent identify based on a clear prevalent identify not being included within the group, wherein the second categorized records comprise one or more additional categories than the first categorized records; and

modifying, by the one or more computing devices, the data records to match the most prevalent identity within the group.

2 . The method of claim 1 , further comprising normalizing, by the one or more computing devices, the data records to be consistent with a predetermined format.

3 . The method of claim 1 , wherein the comparing is performed using regular expression matching or fuzzy matching.

4 . The method of claim 3 , wherein the regular expression matching determines whether two values match when they both satisfy the same regular expression.

5 . The method of claim 3 , wherein the fuzzy matching determines whether two values match when two strings match a pattern approximately.

6 . The method of claim 1 , further comprising:

assigning, by the one or more computing devices, a weight to each of the predetermined categories; and

wherein the similarity score is determined based on the weight for each of the predetermined categories.

7 . The method of claim 1 , wherein modifying the data records to match the most prevalent identity within the group comprises standardizing, by the one or more computing devices, each individual's demographic information within the group to match the most prevalent identity within the group.

8 . A non-transitory computer readable medium having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations for linking ingested data, the operations comprising:

accessing, by one or more computing devices, data records of individuals;

parsing, by the one or more computing devices, the data records to locate each individual's demographic information;

assigning, by the one or more computing devices, the demographic information into predetermined categories;

comparing, by the one or more computing devices, each record against first categorized records using a pair-wise function to determine if they are the same;

training, by the one or more computing devices, a neural network using a training set, wherein the training set is retrieved from an database at a particular time an accuracy of the database was verified;

calculating, using the trained neural network and by the one or more computing devices, a similarity score for each pair of records, wherein the similarity score is a ratio based on how many categories match for each data pair of records;

determining, by the one or more computing devices, whether the similarity score meets or exceeds a similarity score threshold,

based on determining the similarity score meets or exceeds the similarity score threshold, linking, by the one or more computing devices, the data pair in a group;

determining, by the one or more computing devices, a most prevalent identity within the group, comprising:

comparing each record against second categorized records to determine the most prevalent identify based on a clear prevalent identify not being included within the group, wherein the second categorized records comprise one or more additional categories than the first categorized records; and

modifying, by the one or more computing devices, the data records to match the most prevalent identity within the group.

9 . The non-transitory computer readable medium of claim 8 , wherein the operations further comprise normalizing, by the one or more computing devices, the data records to be consistent with a predetermined format.

10 . The non-transitory computer readable medium of claim 8 , wherein the comparing is performed using regular expression matching or fuzzy matching.

11 . The non-transitory computer readable medium of claim 10 , wherein the regular expression matching determines whether two values match when they both satisfy the same regular expression.

12 . The non-transitory computer readable medium of claim 10 , wherein the fuzzy matching determines whether two values match when two strings match a pattern approximately.

13 . The non-transitory computer readable medium of claim 8 , wherein the operations further comprise:

assigning, by the one or more computing devices, a weight to each of the predetermined categories; and

wherein the similarity score is determined based on the weight for each of the predetermined categories.

14 . The non-transitory computer readable medium of claim 8 , wherein modifying the data records to match the most prevalent identity within the group comprises standardizing, by the one or more computing devices, each individual's demographic information within the group to match the most prevalent identity within the group.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2023
From: LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 064968/0455 →
Continuity (2)
Continuation 15948604 · Apr 9, 2018
Related Publication 20230409966A1 · Dec 21, 2023
References Cited (15)
US 9922290B2 · Thomas et al. · 2018 [cited by applicant]
US 20090157573A1 · Anderson et al. · 2009 [cited by applicant]
US 20110307422A1 · Drucker et al. · 2011 [cited by applicant]
US 20140279739A1 · Elkington et al. · 2014 [cited by applicant]
US 20150161210A1 · Cook et al. · 2015 [cited by applicant]
US 20150317563A1 · Baldini Soares et al. · 2015 [cited by applicant]
US 20150379430A1 · Dirac et al. · 2015 [cited by applicant]
US 20170192967A1 · Gilder · 2017 [cited by examiner]
US 20170293666A1 · Ragavan et al. · 2017 [cited by applicant]
US 20170330099A1 · De Vial · 2017 [cited by applicant]
US 20180366114A1 · Anbazhagan · 2018 [cited by examiner]
US 20190138946A1 · Asher et al. · 2019 [cited by applicant]
US 20190311299A1 · Lindner · 2019 [cited by applicant]
WO 2018046378A1 · 2018 [cited by applicant]
Aggarwal, “Data Classification Algorithms and Applications,” 2014, pp. ix-xxii, 1-31, 49, 158, 175, 176, 300. https://doi.org/10.1201/b17320. [cited by applicant]