IP Library › Granted Patent US 10,671,884
Granted Patent B2
US 10,671,884 · App. 16/503,428 · Granted Jun 2, 2020

Systems and methods to improve data clustering using a meta-clustering model

Inventors: Austin Walters (Savoy, IL); Jeremy Goodsitt (Champaign, IL); Anh Truong (Champaign, IL); Reza Farivar (Champaign, IL)
Assignee: Capital One Services, LLC
G06K9/6218G06F17/15G06K9/6267G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,671,884
App. No.
16/503,428
Filed
Jul 3, 2019
Granted
Jun 2, 2020
Kind
B2
Art Unit
2664
USPC
382/225
Abstract

Systems and methods for clustering data are disclosed. For example, a system may include one or more memory units storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data from a client device and generating preliminary clustered data based on the received data, using a plurality of embedding network layers. The operations may include generating a data map based on the preliminary clustered data using a meta-clustering model. The operations may include determining a number of clusters based on the data map using the meta-clustering model and generating final clustered data based on the number of clusters using the meta-clustering model. The operations may include and transmitting the final clustered data to the client device.

Claims (83)

1. A system for clustering data, comprising:

one or more memory units storing instructions; and

one or more processors configured to execute the instructions to perform

operations comprising:

receiving data from a client device;

generating, using a plurality of embedding network layers, preliminary clustered-data based on the received data;

generating, using a meta-clustering model, a data map based on the preliminary clustered-data;

determining, using the meta-clustering model, a number of clusters based on the data map;

generating, using the meta-clustering model, final clustered-data based on the number of clusters; and

transmitting the final clustered-data to the client device.

2. The system of claim 1 , wherein:

the operations further comprise:

generating an updated embedding-network by training an embedding network layer of the embedding network layers, the training being based on the number of data clusters; and

generating, using the updated embedding-network, updated clustered-data based on the received data; and

wherein generating final clustered-data is based on preliminary clustered-data.

3. The system of claim 1 , wherein:

the operations further comprise:

transmitting clustered data samples to the client device, the clustered data samples being based on the clustered data; and

receiving tags associated with the clustered data samples; and

wherein determining the number of clusters is based on the tags.

4. The system of claim 3 , wherein:

the operations further comprise generating tags associated with the clustered data samples using the meta-clustering model; and

determining the number of clusters is based on the tags.

5. The system of claim 1 , wherein generating preliminary clustered-data comprises:

receiving training data;

generating an embedding network layer of the plurality of embedding network layers;

training the embedding network layer to classify training data; and

training the embedding network layer to cluster training data.

6. The system of claim 1 , wherein generating preliminary clustered-data comprises repeating steps until a performance criterion is satisfied, the steps comprising:

adding a trained embedding network layer to the plurality of embedding network layers;

generating clustered data using the added embedding-network;

tagging the clustered data; and

determining whether a performance criterion of the plurality of embedding-networks is satisfied.

7. The system of claim 1 , the operations further comprising:

generating the meta-clustering model;

generating, using the meta-clustering model, encoded data based on the clustered data by reducing the dimensionality of the clustered data;

generating a data map based on the encoded data; and

training the meta-clustering model to determine a number of clusters based on the data map and a performance criterion.

8. The system of claim 7 , wherein training the meta-clustering model comprises iteratively repeating steps until the performance criterion is satisfied, the steps comprising:

determining, using the meta-clustering model, a number of clusters based on the data map;

generating one or more updated embedding-networks based on the number of clusters;

generating updated clustered-data using the one or more updated embedding-networks;

updating the meta-clustering model;

generating, using the updated meta-clustering model, updated encoded-data based on the updated clustered-data by reducing the dimensionality of the updated clustered-data;

generating an updated data-map based on the updated encoded-data; and

determining whether the performance criterion is satisfied based on the updated data-map.

9. The system of claim 1 , wherein the data comprises image data and at least one of the embedding network layers comprises a convolutional neural network.

10. The system of claim 1 , wherein the data comprises text data and at least one of the embedding network layers comprises a language representation model.

11. The system of claim 1 , wherein the embedding network layers comprise at least one of a Bidirectional Encoder Representations from Transformers (BERT) model or an Embeddings from Language Models (ELMo) representation model.

12. The system of claim 1 , wherein the meta-clustering model is a deep learning model.

13. The system of claim 1 , the operations further comprising at least one of transmitting the data map to the client device or transmitting the meta-clustering model to the client device.

14. The system of claim 1 , wherein the number of clusters is larger than a maximum count of clusters in the preliminary clustered-data.

15. The system of claim 1 , wherein generating preliminary clustered-data based on the received data comprises performing a binary classification.

16. The system of claim 1 , wherein generating preliminary clustered-data based on the received data comprises performing one-hot encoding.

17. The system of claim 1 , wherein:

the operations further comprise generating, using the meta-clustering model, encoded data based on the preliminary clustered-data; and

generating the data map is based on the encoded data.

18. The system of claim 17 , wherein generating encoded data comprises performing a method of principal component analysis.

19. A method for clustering data, comprising:

receiving data from a client device;

generating, using a plurality of embedding network layers, preliminary clustered-data based on the received data;

generating, using a meta-clustering model, a data map based on the preliminary clustered-data;

determining, using the meta-clustering model, a number of clusters based on the data map;

generating, using the meta-clustering model, final clustered-data based on the number of clusters; and

transmitting the final clustered-data to the client device.

20. A system for clustering data comprising:

one or more memory units storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

receiving data from a client device;

generating preliminary clustered-data using a plurality of embedding network layers by repeating steps until a performance criterion is satisfied, the steps comprising:

adding a trained embedding network layer;

generating clustered data using the added embedding network layer;

tagging the clustered data; and

determining whether a performance criterion of the plurality of embedding network layers is satisfied;

generating a meta-clustering model comprising a neural network model;

generating, using the meta-clustering model, encoded data based on the clustered data by reducing the dimensionality of the clustered data;

generating a data map based on the encoded data;

training the meta-clustering model to determine a number of clusters based on the data map and a performance criterion;

transmitting clustered data samples to the client device, the clustered data samples being based on the clustered data; and

receiving tags associated with the clustered data samples;

determining, using the meta-clustering model, a number of clusters based on the data map and the tags;

generating, using the meta-clustering model, final clustered-data based on the number of clusters; and

transmitting the final clustered-data to the client device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2019
From: WALTERS, AUSTIN; GOODSITT, JEREMY; TRUONG, ANH; FARIVAR, REZA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 049786/0315 →
Continuity (2)
Provisional Application 62694968 · Jul 6, 2018
Related Publication 20200012886A1 · Jan 9, 2020
Cited By (1)
US 12,347,171