IP Library › Granted Patent US 12,248,756
Granted Patent B2
US 12,248,756 · App. 17/758,276 · Granted Mar 11, 2025

Creating predictor variables for prediction models from unstructured data using natural language processing

Inventors: Howard Hugh Hamilton (Atlanta, GA); Terry Woodford (Kennesaw, GA)
Assignee: Equifax Inc.
G06F40/40G06F40/284G06F40/295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,756
App. No.
17/758,276
Filed
Jun 30, 2022
Granted
Mar 11, 2025
Kind
B2
Examiner
HE, JIALONG
Art Unit
2659
USPC
704/9
Abstract

Systems and methods for creating predictor variables from unstructured data for prediction models are provided. A variable creation application receives unstructured data and processing the unstructured data to generate processed data. Based on the processed data, the variable creation application generates an attribute pool that contains multiple predictor variables generated by applying natural language processing (NLP) procedures on the processed data. The variable creation application further executes a prediction model on at least the predictor variables in the attribute pool to generate a prediction result. Based on the prediction result, the variable creation application evaluates the predictive power of each of the predictor variables and retains predictor variables that are predictive as input predictor variables for the prediction model.

Claims (81)

1. A computer-implemented method in which one or more processing devices performs operations comprising:

accessing unstructured data;

determining a natural language processing (NLP) configuration;

generating a first set of NLP procedures based on the natural language processing configuration;

generating a plurality of predictor variables by applying the first set of NLP procedures on the unstructured data;

generating an attribute pool that comprises the generated plurality of predictor variables;

executing a machine-learning prediction model on at least the plurality of predictor variables in the attribute pool to generate a prediction result;

evaluating a predictive power of each of the plurality of predictor variables based on the prediction result;

updating the NLP configuration;

generating one or more additional NLP procedures not present in the first set of NLP procedures;

generating one or more additional predictor variables by applying the one or more additional NLP procedures;

adding to the attribute pool at least one of the generated one or more additional predictor variables containing a predictive power higher than the predictive power of at least one of the plurality of predictor variables;

retaining at least one predictor variable from the predictor variables within the attribute pool that is predictive as an input predictor variable for the machine-learning prediction model;

training the machine-learning prediction model trained using predictor variables comprising the at least one predictor variable; and

transmitting, to a remote computing device, a risk indicator for a target entity generated by the trained machine-learning prediction model, wherein the risk indicator is usable for controlling access to one or more interactive computing environments by the target entity.

2. The computer-implemented method of claim 1 , further comprising, prior to applying any NLP procedure on the unstructured data, processing the unstructured data by applying one or more of:

a context-independent normalization process;

a context-dependent normalization process;

a tokenization process;

a stop word filtering process; or

a stemming and lemmatization process.

3. The computer-implemented method of claim 1 , wherein the one or more NLP procedures used to generate any predictor variable comprises at least one of a word embedding procedure, a bag-of-word procedure, a named entity recognition procedure, or an information extraction procedure.

4. The computer-implemented method of claim 1 , wherein the operations further comprise:

executing the machine-learning prediction model on the additional predictor variables in the attribute pool to generate additional prediction results.

5. The computer-implemented method of claim 1 , wherein a predictor variable is predictive if the predictive power of the predictor variable satisfies a criterion for predictiveness and is otherwise not predictive.

6. The computer-implemented method of claim 1 , wherein the predictive power of a predictor variable is determined by calculating statistics based on prediction results, the statistics comprising one or more of statistical significance, Kolmogorov-Smirnov (KS) statistics, or Gini statistics.

7. The computer-implemented method of claim 6 , wherein a predictor variable is predictive if the calculated statistics of the predictor variable satisfies a criterion for predictiveness determined by a threshold value of the statistics.

8. A system for generating predictor variables from unstructured data, the system comprising:

one or more processing device; and

one or more non-transitory computer-readable medium communicatively coupled to the one or more processing device, wherein the one or more processing devices are configured to execute program code stored in the non-transitory computer-readable medium and thereby perform operations comprising:

processing the unstructured data to generate processed data;

determining a natural language processing (NLP) configuration;

generating a first set of NLP procedures based on the natural language processing configuration;

generating a plurality of predictor variables by applying the first set of NLP procedures on the unstructured data;

generating an attribute pool that comprises the generated plurality of predictor variables;

executing a prediction model on at least the plurality of predictor variables in the attribute pool to generate a prediction result;

evaluating a predictive power of each of the plurality of predictor variables based on the prediction result;

updating the NLP configuration;

generating one or more additional NLP procedures not present in the first set of NLP procedures;

generating one or more additional predictor variables by applying the one or more additional NLP procedures;

adding to the attribute pool at least one of the generated one or more additional predictor variables; and

retaining at least one predictor variable from the predictor variables within the attribute pool as an input predictor variable for the prediction model.

9. The system of claim 8 , wherein processing the unstructured data comprises applying one or more of:

a context-independent normalization process;

a context-dependent normalization process;

a tokenization process;

a stop word filtering process; or

a stemming and lemmatization process.

10. The system of claim 8 , wherein the one or more NLP procedures used to generate any predictor variable comprises at least one of a word embedding procedure, a bag-of-word procedure, a named entity recognition procedure, or an information extraction procedure.

11. The system of claim 8 , wherein the operations further comprise:

training a machine-learning prediction model trained using predictor variables comprising the at least one predictor variable; and

transmitting, to a remote computing device, a risk indicator for a target entity generated by the trained machine-learning prediction model, wherein the risk indicator is usable for controlling access to one or more interactive computing environments by the target entity.

12. The system of claim 11 , wherein a predictor variable is predictive if the predictive power of the predictor variable satisfies a criterion for predictiveness and is otherwise not predictive.

13. The system of claim 8 , wherein the predictive power of a predictor variable is determined by calculating statistics based on prediction results, the statistics comprising one or more of statistical significance, Kolmogorov-Smirnov (KS) statistics, or Gini statistics.

14. A non-transitory computer-readable medium having instructions stored thereon that are executable by a processor to causing a computing device to perform operations, the operations comprising:

accessing unstructured data;

processing the unstructured data to generate processed data;

determining a natural language processing (NLP) configuration;

generating a first set of NLP procedures based on the natural language processing configuration;

generating a plurality of predictor variables by applying the first set of NLP procedures on the unstructured data;

generating an attribute pool that comprises the generated plurality of predictor variables;

executing a prediction model on at least the plurality of predictor variables in the attribute pool to generate a prediction result;

evaluating a predictive power of each of the plurality of predictor variables based on the prediction result;

updating the NLP configuration;

generating one or more additional NLP procedures not present in the first set of NLP procedures;

generating one or more additional predictor variables by applying the one or more additional NLP procedures;

adding to the attribute pool at least one of the generated one or more additional predictor variables containing a predictive power higher than the predictive power of at least one of the plurality of predictor variables; and

retaining at least one predictor variable from the predictor variables within the attribute pool that is predictive as an input predictor variable for the prediction model.

15. The non-transitory computer-readable medium of claim 14 , wherein processing the unstructured data comprises applying one or more of:

a context-independent normalization process;

a context-dependent normalization process;

a tokenization process;

a stop word filtering process; or

a stemming and lemmatization process.

16. The non-transitory computer-readable medium of claim 14 , wherein the one or more NLP procedures used to generate the plurality of predictor variables comprises at least one of a word embedding procedure, a bag-of-word procedure, a named entity recognition procedure, or an information extraction procedure.

17. The non-transitory computer-readable medium of claim 16 , wherein the operations further comprise:

training a machine-learning prediction model trained using predictor variables comprising the at least one predictor variable; and

transmitting, to a remote computing device, a risk indicator for a target entity generated by the trained machine-learning prediction model, wherein the risk indicator is usable for controlling access to one or more interactive computing environments by the target entity.

18. The non-transitory computer-readable medium of claim 17 , wherein a predictor variable is predictive if the predictive power of the predictor variable satisfies a criterion for predictiveness and is otherwise not predictive.

19. The non-transitory computer-readable medium of claim 14 , wherein the predictive power of a predictor variable is determined by calculating statistics based on prediction results, the statistics comprising one or more of statistical significance, Kolmogorov-Smirnov (KS) statistics, or Gini statistics.

20. The non-transitory computer-readable medium of claim 19 , wherein a predictor variable is predictive if the calculated statistics of the predictor variable satisfies a criterion for predictiveness determined by a threshold value of the statistics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2024
From: WOODFORD, TERRY; HAMILTON, HOWARD HUGH
To: EQUIFAX INC.
Reel/Frame 067886/0477 →
Continuity (2)
Provisional Application 62955100 · Dec 30, 2019
Related Publication 20230023630A1 · Jan 26, 2023
References Cited (14)
US 20140316768A1 · Khandekar · 2014 [cited by examiner]
US 20180025273A1 · Jordan · 2018 [cited by examiner]
US 20180075175A1 · Chang · 2018 [cited by examiner]
US 20190114370A1 · Cerino · 2019 [cited by examiner]
US 20190205771A1 · Lin et al. · 2019 [cited by applicant]
US 20190304582A1 · Blumenthal et al. · 2019 [cited by applicant]
US 20190340526A1 · Turner · 2019 [cited by examiner]
WO 2016160539A1 · 2016 [cited by applicant]
WO 2018084867A1 · 2018 [cited by applicant]
Bach, M. Pejić, et al. “Selection of variables for credit risk data mining models: preliminary research.” 2017 40th International Convention on Information and Communication Technology, Electronics and Microelectronics … [cited by examiner]
Gahlaut, Archana, and Prince Kumar Singh. “Prediction analysis of risky credit using Data mining classification models.” 2017 8th international conference on computing, communication and networking technologies (ICCCNT)… [cited by examiner]
PCT/US2020/067185, “International Preliminary Report on Patentability”, Jul. 14, 2022, 9 pages. [cited by applicant]
PCT/US2020/067185, “International Search Report and Written Opinion”, Mar. 23, 2021, 13 pages. [cited by applicant]
Canadian Application No. CA3,163,408, Office Action, Mailed On Jan. 16, 2024, 6 pages. [cited by applicant]