IP Library › Granted Patent US 12,737,686
Granted Patent B2
US 12,737,686 · App. 18/335,124 · Granted Sep 15, 2026

Updating attribute data structures to indicate trends in attribute data provided to automated modelling systems

Inventors: Jeffrey Q. Ouyang (Atlanta, GA); Vickey Chang (Suwanee, GA); Rupesh Patel (Atlanta, GA); Trevis J. Litherland (Alpharetta, GA)
Assignee: EQUIFAX INC.
G06N20/00G06F16/285G06F17/18G06N7/01G06Q10/063G06Q40/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,686
App. No.
18/335,124
Granted
Sep 15, 2026
Kind
B2
Abstract

Certain aspects involve updating data structures to indicate relationships between attribute trends and response variables used for training automated modeling systems. For example, a data structure stores data for training an automated modeling algorithm. The training data includes attribute values for multiple entities over a time period. A computing system generates, for each entity, at least one trend attribute that is a function of a respective time series of attribute values. The computing system modifies the data structure to include the generated trend attributes and updates the training data to include trend attribute values for the trend attributes. The computing system trains the automated modeling algorithm with the trend attribute values from the data structure. In some aspects, trend attributes are generated by applying a frequency transform to a time series of attribute values and selecting, as trend attributes, some of the coefficients generated by the frequency transform.

Claims (90)

1 . A server system comprising:

a non-transitory computer-readable medium storing a data structure, the data structure comprising training data for training an automated modeling algorithm, the training data comprising data items for a set of attribute values for multiple entities over a time period, the training data mapping a respective attribute value for each of a plurality of attributes to a respective time value within the time period, and an entity identifier for a respective entity associated with the respective attribute value; and

a processing device communicatively coupled to the non-transitory computer-readable medium,

wherein the processing device is configured for performing operations comprising:

generating, for each entity, at least one trend attribute that is a function of a respective time series of attribute values from the set of attribute values, the at least one trend attribute indicating a trend in the time series of attribute values, wherein generating the at least one trend attribute comprises:

(a) generating multiple cluster series, wherein each cluster series comprises clusters respectively associated with intervals in the time period, and

(b) computing, for the multiple cluster series, respective behavioral attribute values, wherein each behavioral attribute value is computed as a function of the respective cluster series;

modifying the data structure to include the at least one trend attribute, wherein modifying the data structure includes storing entries mapping the respective entity identifiers to a respective plurality of trend attribute values;

updating the training data by performing, for at least some of the entities in the training data, operations comprising:

identifying, for each entity, a respective cluster series having a respective behavioral attribute value that is similar to a respective behavior of a respective time series of attributes values for the entity,

assigning a cluster membership to the entity based on the respective behavioral attribute value being similar to the respective behavior of the respective time series of attributes values for the entity, and

selecting, for the entity, an identifier of the cluster membership as a trend attribute value for the entity; and

training the automated modeling algorithm with the trend attribute values from the data structure, including providing the training data to the automated modeling algorithm to train the automated modeling algorithm to teach the automated modeling algorithm to predict a risk outcome.

2 . The server system of claim 1 , wherein the processing device is configured for generating the at least one trend attribute by performing, for each entity, additional operations comprising:

identifying a respective subset of attribute values for at least one attribute based on the attribute values being associated with the entity;

applying a frequency transform to the respective subset of attribute values; and

selecting, as the at least one trend attribute, at least one coefficient generated by the applied frequency transform.

3 . The server system of claim 1 , wherein the function of the respective time series uses changes in the respective time series over the time period to compute a trend attribute value.

4 . The server system of claim 1 , wherein the updated training data comprises first data items having a first trend attribute value and second data items having a second trend attribute value;

wherein the processing device is configured for training the automated modeling algorithm by executing segmentation logic based on the trend attribute values, wherein executing the segmentation logic comprises:

applying, based on the first data items having the first trend attribute value, a first modeling function to the first data items, and

applying, based on the second data items having the second trend attribute value, a second modeling function to the second data items.

5 . The server system of claim 1 , wherein the processing device is configured for grouping the respective subset of attribute values into the respective cluster by performing operations comprising:

applying, for each entity, a frequency transform to attribute values associated with the entity; and

grouping the respective subset of attribute values into the respective cluster based on at least one coefficient generated by the applied frequency transform.

6 . The server system of claim 1 , wherein the processing device is configured for grouping the respective subset of attribute values into the respective cluster by performing operations comprising:

identifying a first time series of attribute values for a first attribute and a second time series of attribute values for a second attribute;

performing a principal component analysis on the first time series and the second time series;

outputting a principal component data series from the principal component analysis; and

grouping the principal component data series into the clusters.

7 . A method comprising:

accessing a data structure having training data for an automated modeling algorithm, the training data comprising data items with attribute values for multiple entities over a time period, the training data mapping a respective attribute value for each of a plurality of attributes to a respective time value within the time period, and an entity identifier for a respective entity associated with the respective attribute value;

generating, with a processing device, at least one trend attribute that is a function of a time series of attribute values, the at least one trend attribute indicating a trend in the time series of attribute values, wherein generating the at least one trend attribute comprises:

applying, for each entity, a frequency transform to a respective subset of attribute values for the at least one attribute based on the respective subset of attribute values being associated with the entity, and

selecting, as the at least one trend attribute, at least one coefficient generated by the applied frequency transform;

modifying, with the processing device, the data structure to include the at least one trend attribute, wherein modifying the data structure includes storing entries mapping the respective entity identifiers to a respective plurality of trend attribute values;

updating, with the processing device, the training data to include trend attribute values for the at least one trend attribute, the trend attribute values including values of the at least one coefficient, wherein updating the training data comprises performing, for at least some of the entities in the training data, operations comprising:

identifying, for each entity, a respective cluster series having a respective behavioral attribute value that is similar to a respective behavior of a respective time series of attributes values for the entity,

assigning a cluster membership to the entity based on the respective behavioral attribute value being similar to the respective behavior of the respective time series of attributes values for the entity, and

selecting, for the entity, an identifier of the cluster membership as a trend attribute value for the entity;

training, with the processing device, the automated modeling algorithm with the trend attribute values and one or more attribute values from the data structure, including providing the training data to the automated modeling algorithm to train the automated modeling algorithm to teach the automated modeling algorithm to predict a risk outcome; and

outputting, with the processing device, the trend attribute values from the data structure to a computing system that executes the automated modeling algorithm.

8 . The method of claim 7 , further comprising generating at least one additional trend attribute by performing operations comprising:

identifying intervals in the time period; and

generating multiple cluster series, wherein each cluster series comprises clusters respectively associated with the intervals, wherein the processing device is configured for generating each cluster series by performing operations comprising:

computing, for the multiple cluster series, respective additional trend attribute values, wherein each additional trend attribute value is computed as a function of the respective cluster series.

9 . The method of claim 8 , wherein the updated training data comprises first data items having a first trend attribute value and second data items having a second trend attribute value;

wherein training the automated modeling algorithm comprises executing segmentation logic based on the additional trend attribute values, wherein executing the segmentation logic comprises:

applying, based on the first data items having the first trend attribute value, a first modeling function to the first data items, and

applying, based on the second data items having the second trend attribute value, a second modeling function to the second data items.

10 . The method of claim 7 , wherein the function of the respective time series uses changes in the respective time series over the time period to compute a trend attribute value, wherein the at least one trend attribute comprises at least one of:

a statistical attribute,

a duration attribute computed based on peaks and valleys in the respective time series, or a depression/recovery attribute computed based on rates of change between the peaks and the valleys in the respective time series,

a skewness of a probability distribution of the respective time series, or

a kurtosis of a probability distribution of the respective time series.

11 . The method of claim 7 , wherein the respective subset of attribute values are grouped into the respective cluster based on the at least one coefficient generated by the frequency transform.

12 . The method of claim 7 , wherein grouping the respective subset of attribute values into the respective cluster comprises:

identifying a first time series of attribute values for a first attribute and a second time series of attribute values for a second attribute;

performing a principal component analysis on the first time series and the second time series;

outputting a principal component data series from the principal component analysis; and

grouping the principal component data series into the clusters.

13 . A non-transitory computer-readable medium having program code that is executable by a processing device to cause a computing device to perform operations, the operations comprising:

accessing a data structure comprising training data for training an automated modeling algorithm, the training data comprising data items with attribute values for multiple entities over a time period, the training data mapping a respective attribute value for each of a plurality of attributes to a respective time value within the time period, and an entity identifier for a respective entity associated with the respective attribute value;

generating at least one trend attribute that is a function of a time series of attribute values, the at least one trend attribute indicating a trend in the time series of attribute values, wherein generating the at least one trend attribute comprises:

(a) generating multiple cluster series, wherein each cluster series comprises clusters respectively associated with intervals in the time period, and

(b) computing, for the multiple cluster series, respective behavioral attribute values, wherein each behavioral attribute value is computed as a function of the respective cluster series;

modifying the data structure to include the at least one trend attribute, wherein modifying the data structure includes storing entries mapping the respective entity identifiers to a respective plurality of trend attribute values;

updating the training data to include trend attribute values for the at least one trend attribute, wherein updating the training data comprises performing, for at least some of the entities in the training data comprises:

identifying, for each entity, a respective cluster series having a respective behavioral attribute value that is similar to a respective behavior of a respective time series of attributes values for the entity,

assigning a cluster membership to the entity based on the respective behavioral attribute value being similar to the respective behavior of the respective time series of attributes values for the entity, and

selecting, for the entity, an identifier of the cluster membership as a trend attribute value for the entity; and

training the automated modeling algorithm with the trend attribute values and one or more attribute values from the data structure, including providing the training data to the automated modeling algorithm to train the automated modeling algorithm to teach the automated modeling algorithm to predict a risk outcome.

14 . The non-transitory computer-readable medium of claim 13 , wherein generating the at least one trend attribute comprises performing, for each entity, operations comprising:

identifying a respective subset of attribute values for the at least one attribute based on a selected subset of attribute values being associated with the entity;

applying a frequency transform to the respective subset of attribute values; and

selecting, as the at least one trend attribute, at least one coefficient generated by the applied frequency transform.

15 . The non-transitory computer-readable medium of claim 13 , wherein the function of the respective time series uses changes in the respective time series over the time period to compute a trend attribute value, wherein the at least one trend attribute comprises at least one of:

a statistical attribute,

a duration attribute computed based on peaks and valleys in the respective time series, or a depression/recovery attribute computed based on rates of change between the peaks and the valleys in the respective time series,

a skewness of a probability distribution of the respective time series, or

a kurtosis of a probability distribution of the respective time series.

16 . The non-transitory computer-readable medium of claim 13 , wherein the updated training data comprises first data items having a first trend attribute value and second data items having a second trend attribute value;

wherein training the automated modeling algorithm comprises executing segmentation logic based on the trend attribute values, wherein executing the segmentation logic comprises:

applying, based on the first data items having the first trend attribute value, a first modeling function to the first data items, and

applying, based on the second data items having the second trend attribute value, a second modeling function to the second data items.

17 . The non-transitory computer-readable medium of claim 13 , wherein grouping the respective subset of attribute values into the respective cluster comprises:

identifying a first time series of attribute values for a first attribute and a second time series of attribute values for a second attribute;

performing a principal component analysis on the first time series and the second time series;

outputting a principal component data series from the principal component analysis; and

grouping the principal component data series into the clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2024
From: OUYANG, JEFFREY Q.; CHANG, VICKEY; PATEL, RUPESH; LITHERLAND, TREVIS J.
To: EQUIFAX INC.
Reel/Frame 066282/0268 →
Continuity (3)
Continuation 15761864 · Sep 21, 2016
Provisional Application 62221360 · Sep 21, 2015
Related Publication 20230325724A1 · Oct 12, 2023
References Cited (42)
US 10572945B1 · McNair · 2020 [cited by examiner]
US 11715029B2 · Ouyang et al. · 2023 [cited by applicant]
US 20030093352A1 · Muralidhar et al. · 2003 [cited by applicant]
US 20110040723A1 · O'Donnell et al. · 2011 [cited by applicant]
US 20110246385A1 · Laxmanan · 2011 [cited by examiner]
US 20110251870A1 · Tavares et al. · 2011 [cited by applicant]
US 20130132390A1 · Ghosh · 2013 [cited by examiner]
US 20150170056A1 · Breckenridge et al. · 2015 [cited by applicant]
WO 2017053347A1 · 2017 [cited by applicant]
Birvinskas, D. et al, EEG Dataset Reduction and Feature Extraction using Discrete Cosine Transform [online], 2012 [retrieved on Jun. 24, 2025]. Retrieved from Internet: <https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&a… [cited by examiner]
Gulutzan, P., MySQL Stored Procedures, [retrieved on Jun. 24, 2025]. Retrieved from Internet:<chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/http://dtucker.cs.edinboro.edu/CSCI313/mysql-stored-procedures.pdf> (Year… [cited by examiner]
Tsai, et al, Credit rating by hybrid machine learning techniques, [retrieved Jun. 24, 2025]. Retrieved from Internet:<https://www.sciencedirect.com/science/article/pii/S1568494609001215> (Year: 2009). [cited by examiner]
Lai, et al, Evolving and clustering fuzzy decision tree for financial time series data forecasting, Retrieved from Internet:<https://www.sciencedirect.com/science/article/pii/S0957417408001474> (Year: 2009). [cited by examiner]
Canadian Application No. CA2999276 , Notice of Allowance, Mailed On Mar. 14, 2025, 1 page. [cited by applicant]
European No. EP16849456.5, “Intention to Grant”, mailed May 30, 2025, 9 pages. [cited by applicant]
Canadian Application No. CA2,999,276 , “Office Action”, mailed Sep. 5, 2023, 7 pages. [cited by applicant]
“BCI Competition Datasets”, Available Online at: http://bbci.de/competition, 2012, 3 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Advisory Action”, Mar. 21, 2022, 4 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Final Office Action”, Nov. 9, 2021, 20 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Final Office Action”, Oct. 26, 2022, 26 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Non-Final Office Action”, Apr. 26, 2022, 29 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Non-Final Office Action”, May 19, 2021, 21 pages. [cited by applicant]
U.S. Appl. No. 15/761,864, “Notice of Allowance”, Mar. 14, 2023, 13 pages. [cited by applicant]
Aghabozorgi, et al., “Time-Series Clustering—A Decade Review”, Information Systems, vol. 53, May 6, 2015, pp. 16-38. [cited by applicant]
Australian Patent Application No. 2016328959, “First Examination Report”, Feb. 9, 2021, 5 pages. [cited by applicant]
Australian Patent Application No. 2016328959, “Notice of Acceptance”, Jun. 16, 2021, 3 pages. [cited by applicant]
Australian Patent Application No. 2021232839, “First Examination Report”, Oct. 24, 2022, 4 pages. [cited by applicant]
Australian Patent Application No. 2021232839, “Notice of Acceptance”, Jun. 19, 2023, 3 pages. [cited by applicant]
Birvinskas, et al., “EEG Dataset Reduction and Feature Extraction Using Discrete Cosine Transform”, Computer Modeling and Simulation (EMS), Sixth Uksim/Amss European Symposium on IEEE, Nov. 14, 2012, pp. 199-204. [cited by applicant]
Canadian Patent Application No. 2,999,276, “Office Action”, Sep. 26, 2022, 4 pages. [cited by applicant]
European Patent Application No. 16849456.5, “Extended European Search Report”, May 28, 2019, 9 pages. [cited by applicant]
European Patent Application No. 16849456.5, “Office Action”, Mar. 30, 2023, 5 pages. [cited by applicant]
European Patent Application No. 16849456.5, “Office Action”, Oct. 8, 2021, 8 pages. [cited by applicant]
Gulutzan, “MySQL Stored Procedures”, 2006, 109 pages. [cited by applicant]
Indian Patent Application No. IN201817010214, “First Examination Report”, Feb. 26, 2021, 7 pages. [cited by applicant]
Khandani, et al., “Consumer Credit-Risk Models via Machine-Learning Algorithms”, Journal of Banking & Finance, vol. 34, No. 11, Available Online at: https://www.sciencedirect.com/science/article/pii/S0378426610002372, N… [cited by applicant]
Liao, “Clustering of Time Series Data—A Survey”, Pattern Recognition, vol. 38, No. 11, Nov. 1, 2005, pp. 1857-1874. [cited by applicant]
PCT/US2016/052759, “International Preliminary Report on Patentability”, Apr. 5, 2018, 9 pages. [cited by applicant]
PCT/US2016/052759, “International Search Report and Written Opinion”, Dec. 30, 2016, 12 pages. [cited by applicant]
Shayegan, et al., “A New Dataset Size Reduction Approach for PCA-Based Classification in OCR Application”, Mathematical Problems in Engineering, vol. 2014, Apr. 17, 2014, pp. 1-14. [cited by applicant]
Tsai, et al., “Credit Rating by Hybrid Machine Learning Techniques”, Applied Soft Computing, 2009. [cited by applicant]
Wang, et al., “Training Data Selection for Support Vector Machines”, Advances in Natural Computation; [Lecture Notes in Computer Science;Lncs], Springer-Verlag, Berlin/ Heidelberg, Jul. 23, 2005, pp. 554-564. [cited by applicant]