IP Library › Granted Patent US 11,531,875
Granted Patent B2
US 11,531,875 · App. 15/931,369 · Granted Dec 20, 2022

Systems and methods for generating datasets for model retraining

Inventors: Anand Dwivedi (Boston, MA); Hyunsoo Jeong (Boston, MA)
Assignee: NASDAQ, INC.
G06N3/08G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,531,875
App. No.
15/931,369
Granted
Dec 20, 2022
Kind
B2
Abstract

A computer system is provided and programmed to assemble a plurality of synthetic datasets and blend those synthetic datasets into a synthesized dataset. An evaluation is then performed to determine whether an existing model should be associated with the synthesized dataset or a new model should be trained from an existing model using the synthesized dataset.

Claims (71)

1. A computer system comprising:

non-transitory computer readable memory that is configured to store:

a reference model; and

a reference dataset that is associated with the reference model;

a processing system that includes at least one hardware processor, the processing system configured to:

generate a plurality of synthetic datasets that are derived from labeled detection frames;

generate, for each synthetic dataset of the plurality of synthetic datasets, a plurality of feature metrics for a plurality of features from the each synthetic dataset, wherein the feature metrics are generated based on the reference dataset;

use a first neural network to generate, based on the determined plurality of feature metrics, a dataset similarity score for each of the plurality of synthetic datasets with respect to the reference dataset, wherein each of the dataset similarity scores indicates how similar a given synthetic dataset is to the reference dataset;

generate, for each of the plurality of synthetic datasets, a training similarity score by training a neural network architecture of the reference model by using a corresponding synthetic dataset; and

generate a synthesized dataset by combining data from the plurality of synthetic datasets based on the training similarity scores and the dataset similarity scores.

2. The system of claim 1 , wherein the processing system is further configured to:

select features from each of the plurality synthetic datasets that have separability that is greater than a threshold amount,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated based on those features that are selected.

3. The system of claim 1 , wherein the processing system is further configured to:

perform a feature level similarly process between each of the plurality synthetic datasets and the reference dataset,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated based on the performed feature level similarly process.

4. The system of claim 3 , wherein the processing system is further configured to:

calculate a density estimate curve, with respect to the reference dataset, for each feature for each of the plurality synthetic datasets,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated the calculated density estimate curves for the features in the plurality synthetic datasets.

5. The system of claim 4 , wherein the processing system is further configured to:

calculate, for each feature of each of the plurality synthetic datasets, a geometric similarity based on a corresponding calculated density estimate curve.

6. The system of claim 1 , wherein the processing system is further configured to:

perform a sample-level similarity check that includes a homogeneity check and a heterogeneity check, the homogeneity check measuring how similar the same classes are between the reference dataset and one of the plurality of synthetic datasets, the heterogeneity check measuring how dissimilar different classes are within the same synthetic dataset,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated the calculated density estimate curves for the features in the plurality synthetic datasets.

7. The system of claim 1 , wherein the processing system is further configured to:

perform a Model-Agnostic Tensor Homogeneity evaluator process to calculate the plurality of feature metrics.

8. The system of claim 1 , wherein the processing system is further configured to:

test performance of the synthesized dataset against the reference model; and

based on determination that the tested performance of the synthesized dataset is within a threshold amount, store an association between the synthesized dataset and the reference model.

9. The system of claim 8 , wherein the processing system is further configured to:

based on the determination that the tested performance of the synthesized dataset is outside the threshold amount, train a new model by using the synthesized dataset; and

store an association between the synthesized dataset and the new model.

10. A method implemented on a computer system, the method comprising:

storing, to a non-transitory storage medium, a reference model and a reference dataset that is associated with the reference model;

generating a plurality of synthetic datasets that are derived from labeled detection frames;

generating, for each synthetic dataset of the plurality of synthetic datasets, a plurality of feature metrics for a plurality of features from the synthetic dataset, wherein the feature metrics are generated based on comparison to the reference dataset;

using a first neural network to generate, based on the determined plurality of feature metrics, a dataset similarity score for each of the plurality of synthetic datasets with respect to the reference dataset, wherein each of the dataset similarity scores indicates how similar a given synthetic dataset is to the reference dataset;

generating, for each of the plurality of synthetic datasets, a training similarity score by training a neural network architecture of the reference model by using a corresponding synthetic dataset; and

constructing a synthesized dataset by combining data from the plurality of synthetic datasets based on the training similarity scores and the dataset similarity scores.

11. The method of claim 10 , further comprising:

selecting features from each of the plurality synthetic datasets that have separability that is greater than a threshold amount,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated based on those features that are selected.

12. The method of claim 10 , further comprising:

performing a feature level similarly process between each of the plurality synthetic datasets and the reference dataset,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated based on the performed feature level similarly process.

13. The method of claim 10 , further comprising:

calculating a density estimate curve, with respect to the reference dataset, for each feature for each of the plurality synthetic datasets,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated the calculated density estimate curves for the features in the plurality synthetic datasets.

14. The method of claim 10 , further comprising:

calculating, for each feature of each of the plurality synthetic datasets, a geometric similarity based on a corresponding calculated density estimate curve.

15. The method of claim 10 , further comprising:

performing a sample-level similarity check that includes a homogeneity check and a heterogeneity check, the homogeneity check measuring how similar the same classes are between the reference dataset and one of the plurality of synthetic datasets, the heterogeneity check measuring how dissimilar different classes are within the same synthetic dataset,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated the calculated density estimate curves for the features in the plurality synthetic datasets.

16. The method of claim 10 , further comprising:

performing a Model-Agnostic Tensor Homogeneity Evaluator process to calculate the plurality of feature metrics.

17. The method of claim 10 , further comprising:

testing performance of the synthesized dataset against the reference model; and

based on determination that the tested performance of the synthesized dataset is within a threshold amount, storing an association between the synthesized dataset and the reference model.

18. The method of claim 10 , further comprising:

based on the determination that the tested performance of the synthesized dataset is outside the threshold amount, training a new model by using the synthesized dataset; and

storing an association between the synthesized dataset and the new model.

19. A non-transitory computer readable storage medium configured to store computer-executable instructions for use with a computer system, the stored computer-executable instructions comprising instructions that cause the computer system to perform operations comprising:

storing, to a non-transitory storage medium, a reference model and a reference dataset that is associated with the reference model;

generating a plurality of synthetic datasets that are derived from labeled detection frames;

generating, for each synthetic dataset of the plurality of synthetic datasets, a plurality of feature metrics for a plurality of features from the synthetic dataset, wherein the feature metrics are generated based on comparison to the reference dataset;

using a first neural network to generate, based on the determined plurality of feature metrics, a dataset similarity score for each of the plurality of synthetic datasets with respect to the reference dataset, wherein each of the dataset similarity scores indicates how similar a given synthetic dataset is to the reference dataset;

generating, for each of the plurality of synthetic datasets, a training similarity score by training a neural network architecture of the reference model by using a corresponding synthetic dataset; and

constructing a synthesized dataset by combining data from the plurality of synthetic datasets based on the training similarity scores and the dataset similarity scores.

20. The non-transitory computer readable storage medium of claim 19 , wherein the operations further comprise:

selecting features from each of the plurality synthetic datasets that have separability that is greater than a threshold amount,

wherein the plurality of feature metrics for each of the plurality synthetic datasets are generated based on those features that are selected.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2020
From: DWIVEDI, ANAND; JEONG, HYUNSOO
To: NASDAQ, INC.
Reel/Frame 053612/0188 →
Continuity (2)
Provisional Application 62847621 · May 14, 2019
Related Publication 20200364551A1 · Nov 19, 2020