IP Library › Granted Patent US 11,604,896
Granted Patent B2
US 11,604,896 · App. 16/889,363 · Granted Mar 14, 2023

Systems and methods to improve data clustering using a meta-clustering model

Inventors: Austin Walters (Savoy, IL); Jeremy Goodsitt (Champaign, IL); Anh Truong (Champaign, IL); Reza Farivar (Champaign, IL)
Assignee: Capital One Services, LLC
G06F21/6254G06F8/71G06F9/54G06F9/541G06F9/547G06F11/3608G06F11/3628G06F11/3636G06F16/2237G06F16/2264G06F16/248G06F16/2423G06F16/24568G06F16/254G06F16/258G06F16/283G06F16/285G06F16/288G06F16/335G06F16/906G06F16/9038G06F16/90332G06F16/90335G06F16/93G06F17/15G06F17/16G06F17/18G06F21/552G06F21/60G06F21/6245G06F30/20G06F40/117G06F40/166G06F40/20G06K9/6215G06K9/6218G06K9/6227G06K9/6231G06K9/6232G06K9/6253G06K9/6256G06K9/6257G06K9/6262G06K9/6265G06K9/6267G06K9/6269G06K9/6277G06N3/04G06N3/0445G06N3/0454G06N3/08G06N3/088G06N5/00G06N5/02G06N5/04G06N7/00G06N7/005G06N20/00G06Q10/04G06T7/194G06T7/246G06T7/248G06T7/254G06T11/001G06V10/768G06V10/993G06V30/194G06V30/1985H04L63/1416H04L63/1491H04L67/306H04L67/34H04N21/23412H04N21/8153G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,604,896
App. No.
16/889,363
Granted
Mar 14, 2023
Kind
B2
Abstract

Systems and methods for clustering data are disclosed. For example, a system may include one or more memory units storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving data from a client device and generating preliminary clustered data based on the received data, using a plurality of embedding network layers. The operations may include generating a data map based on the preliminary clustered data using a meta-clustering model. The operations may include determining a number of clusters based on the data map using the meta-clustering model and generating final clustered data based on the number of clusters using the meta-clustering model. The operations may include and transmitting the final clustered data to the client device.

Claims (75)

1. A system for clustering data, comprising:

one or more memory units storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

receiving data from a database;

generating, using a plurality of embedding network layers, preliminary data clusters based on the received data;

generating, using a meta-clustering model, final clustered data based on the preliminary data clusters; and

storing the final clustered data.

2. The system of claim 1 , the operations further comprising:

generating an updated embedding network by training an embedding network layer of the embedding network layers; and

generating, using the updated embedding-network, updated clustered data based on the received data.

3. The system of claim 1 , the operations further comprising:

determining, using the meta-clustering model, a number of clusters based on the received data.

4. The system of claim 3 , the operations further comprising:

transmitting clustered data samples to a client device configured to generate a user interface that receives tags from a user, the clustered data samples being based on the clustered data; and

receiving, from the client device, tags associated with the clustered data samples; and wherein determining the number of clusters comprises determining the number of clusters based on the tags.

5. The system of claim 1 , the operations further comprising:

sending, to the client device for display on the user interface, at least one of:

the final clustered data;

a visual representation of the final clustered data;

the number of clusters; or

the meta-clustering model.

6. The system of claim 1 , the operations further comprising:

generating clustered data samples based on the preliminary data clusters; and

generating tags associated with the clustered data samples using the meta-clustering model,

wherein determining the number of clusters is based on the tags.

7. The system of claim 1 , wherein generating preliminary data clusters comprises:

receiving training data;

generating one of the embedding network layers;

training the generated embedding network layer to classify training data; and

training the generated embedding network layer to cluster training data.

8. The system of claim 1 , wherein generating preliminary data clusters comprises repeating generating steps until a performance criterion is satisfied, the generating steps comprising:

adding a trained embedding network layer to the plurality of embedding network layers;

generating clustered data using the added embedding-network layer;

tagging the clustered data; and

determining whether a performance criterion of the embedding networks is satisfied.

9. The system of claim 1 , the operations further comprising:

generating the meta-clustering model;

reducing a dimensionality of the clustered data;

generating, using the meta-clustering model, encoded data based on the reduced-dimensionality clustered data;

generating a data map based on the encoded data; and

training the meta-clustering model to determine a number of clusters based on the data map and a performance criterion.

10. The system of claim 8 , wherein generating, using the meta-clustering model, final clustered data is further based on the number of clusters.

11. The system of claim 8 , wherein training the meta-clustering model comprises iteratively repeating training steps until the performance criterion is satisfied, the training steps comprising:

determining, using the meta-clustering model, a number of clusters based on the data map;

generating an updated embedding network based on the number of clusters;

generating updated clustered-data using the updated embedding network;

updating the meta-clustering model;

reducing a dimensionality of the updated clustered data;

generating, using the updated meta-clustering model, updated encoded data based on the reduced-dimensionality updated clustered data;

generating an updated data map based on the updated encoded data; and

determining whether the performance criterion is satisfied based on the updated data map.

12. The system of claim 1 , wherein the data comprises image data and the embedding network layer comprises a convolutional neural network.

13. The system of claim 1 , wherein the data comprises text data and the embedding network layers comprise a language representation model.

14. The system of claim 1 , wherein the embedding network layers comprise at least one of a Bidirectional Encoder Representations from Transformers (BERT) model or an Embeddings from Language Models (ELMo) representation model.

15. The system of claim 1 , wherein the meta-clustering model comprises a deep learning model.

16. The system of claim 3 , wherein the number of clusters is larger than a maximum count of clusters in the preliminary data clusters.

17. The system of claim 1 , wherein generating preliminary data clusters comprises performing a binary classification.

18. The system of claim 1 , wherein generating preliminary data clusters comprises performing one-hot encoding.

19. A method for clustering data, comprising:

receiving data from a database;

generating, using a plurality of embedding network layers, preliminary data clusters based on the received data;

generating, using a meta-clustering model, final clustered data based on the preliminary data clusters; and

storing the final clustered data.

20. A system for clustering data comprising:

one or more memory units storing instructions; and

one or more processors configured to execute the instructions to perform operations comprising:

generating preliminary data clusters using a plurality of embedding network layers by repeating steps until a performance criterion is satisfied, the steps comprising:

adding a trained embedding network layer;

generating clustered data using the trained embedding network layer;

tagging the clustered data; and

determining whether a performance criterion of the embedding network layers is satisfied;

reducing a dimensionality of the clustered data;

generating, using a meta-clustering model, encoded data based on the reduced-dimensionality clustered data;

generating a data map based on the encoded data; and

generating, using the meta-clustering model, final clustered data based on the data map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2020
From: WALTERS, AUSTIN; GOODSITT, JEREMY; TRUONG, ANH; FARIVAR, REZA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 052803/0399 →
Continuity (3)
Continuation 16503428 · Jul 3, 2019
Provisional Application 62694968 · Jul 6, 2018
Related Publication 20200293427A1 · Sep 17, 2020
Cited By (1)
US 12,438,790