IP Library › Granted Patent US 11,367,142
Granted Patent B1
US 11,367,142 · App. 16/732,156 · Granted Jun 21, 2022

Systems and methods for clustering data to forecast risk and other metrics

Inventors: Wensu Wang (Katy, TX); Chun Wang (Austin, TX); Ji Zang (Sugar Land, TX); Kuikui Gao (Houston, TX); Wanli Cheng (Houston, TX)
Assignee: DatalnfoCom USA, Inc.
G06Q40/08G06F16/2456G06N20/00G06Q10/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,367,142
App. No.
16/732,156
Filed
Dec 31, 2019
Granted
Jun 21, 2022
Kind
B1
Examiner
LEE, WILSON
Art Unit
2152
USPC
707/756
Abstract

Methods and systems are provided for clustering data for use in training models to predict a loss metric for insurance. The historical data is cleaned and aggregated to a quarterly time granularity and a zip code geographic granularity, and features on which to generate the clusters are selected. Clusters are generated using one or more clustering algorithms, evaluated, and the best set of clusters is selected. Clusters may be grouped together based on cluster characteristics. Each cluster or group of clusters is used to train a development model and a forecast model for the loss metric. The accuracy of the development and forecast models is evaluated, and the clustering process is repeated until the accuracy and stability of the development and forecast models is sufficient. At each iteration, historical data may be removed and/or hyperparameters of the clustering algorithms may be adjusted.

Claims (45)

1. A computer-implemented method of clustering historical data, the method comprising:

receiving historical data comprising historical policyholder data, historical policy data, historical claims data, historical environmental data, and historical loss metric data, the historical data further comprising at least one input variable;

processing the historical data, comprising:

joining the historical data;

cleaning the historical data; and

aggregating the historical data to a specified time granularity and specified second granularity;

generating a set of features based on the processed historical data;

inputting the feature set into a plurality of clustering algorithms to generate a plurality of sets of clusters, each clustering algorithm generating a set of clusters based on the feature set, each clustering algorithm comprising one of the following: k-means, k-medians, mean shift, DBSCAN, Gaussian mixture model, hierarchical clustering, affinity propagation, and spectral clustering;

evaluating each of the plurality of sets of clusters for stability; and

selecting the most stable set of clusters.

2. The method of claim 1 , wherein the specified time granularity is quarterly.

3. The method of claim 1 , wherein the specified second granularity is a geographic granularity.

4. The method of claim 1 , wherein generating the set of features comprises:

adding input variables of known importance to the feature set; and

adding at least one newly created variable to the feature set.

5. The method of claim 4 , wherein the newly created variable comprises a smoothed loss metric value.

6. The method of claim 4 , wherein the newly created variable comprises a mean of the loss metric over time.

7. The method of claim 4 , wherein the newly created variable comprises a trend of the loss metric over time.

8. The method of claim 4 , wherein the newly created variable comprises a standard deviation of the loss metric over time.

9. The method of claim 1 , wherein the stability of a set of clusters is determined based on the trend of the loss metric of each cluster.

10. The method of claim 1 , wherein the stability of a set of clusters is determined based on the separation between the loss metrics of each cluster.

11. The method of claim 1 , wherein the stability of a set of clusters is determined based on the exposure coverage of the set of clusters.

12. The method of claim 1 , further comprising:

for each cluster in the set of clusters, training a loss metric model based the historical data comprising the cluster.

13. A method of clustering historical data, the method comprising:

receiving historical data comprising historical policyholder data, historical policy data, historical claims data, historical external data, and historical loss metric data, the historical data further comprising at least one input variable;

segmenting the historical data into a plurality of segments;

processing the historical data, comprising:

joining the historical data;

cleaning the historical data; and

for each segment, aggregating the historical data to a specified time granularity and specified second granularity;

for each segment, generating a set of features based on the processed historical data;

for each segment, inputting the feature set into a plurality of clustering algorithms to generate a plurality of sets of clusters, each clustering algorithm generating a set of clusters based on the feature set, each clustering algorithm comprising one of the following: k-means, k-medians, mean shift, DBSCAN, Gaussian mixture model, hierarchical clustering, affinity propagation, and spectral clustering;

evaluating each of the plurality of sets of clusters for stability; and

for each segment, selecting the most stable set of clusters.

14. The method of claim 13 , further comprising

grouping a cluster from one segment with a cluster from a different segment based on an attribute of the loss metric of the clusters.

15. The method of claim 14 , wherein the attribute relates to the value of the loss metric.

16. The method of claim 13 , wherein the specified time granularity is quarterly.

17. The method of claim 13 , wherein the specified second granularity is a geographic granularity.

18. The method of claim 13 , wherein generating the set of features comprises:

adding input variables of known importance to the feature set; and

adding at least one newly created variable to the feature set.

19. The method of claim 18 , wherein the newly created variable comprises one of the following: a smoothed loss metric value, the mean of the loss metric over time, the trend of the loss metric over time, and the standard deviation of the loss metric over time.

20. The method of claim 13 , wherein the stability of a set of clusters is determined based on one or more of the following: the trend of the loss metric of each cluster, the separation between the loss metrics of each cluster, and the exposure coverage of the set of clusters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2020
From: WANG, CHUN; WANG, WENSU; ZANG, JI; GAO, KUIKUI; CHENG, WANLI
To: DATAINFOCOM USA, INC.
Reel/Frame 051942/0665 →
Continuity (2)
Continuation In Part 16146590 · Sep 28, 2018
Provisional Application 62564468 · Sep 28, 2017
Cited By (5)
US 12,284,071 US 12,462,208 US 12,555,418 US 12,670,480 US 12,694,456