IP Library Granted Patent US 12,046,021
Granted Patent B2
US 12,046,021 · App. 17/306,006 · Granted Jul 23, 2024

Machine learning training dataset optimization

Inventors: Or Shabtay (Tel-Aviv, IL); Eran Shlomo (Zichron Yaakov, IL)
Assignee: DATALOOP LTD.
G06V10/774G06F18/2115G06F18/2148G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,046,021
App. No.
17/306,006
Granted
Jul 23, 2024
Kind
B2
Abstract

A method comprising: receiving a dataset comprising a plurality of data instances; extracting a feature vector representation of each of the data instances in the dataset; choosing a first data instance for adding to a subset of the dataset, wherein the first data instance is removed from the dataset; performing an iterative process comprising: (i) identifying one of the data instances in the dataset which represents a maximal information addition to the subset, based, at least in part, on measuring an information difference parameter between the feature vector representation of the identified data instance and the feature vector representations of all of the data instances in the subset, and (ii) adding the identified data instance to the subset and removing the identified data instance from the dataset, until the information difference parameter is lower than a predetermined threshold; and outputting the subset as a representative subset of the dataset.

Claims (46)

1. A system comprising:

at least one hardware processor; and

a non-transitory computer-readable storage medium having stored thereon program instructions, the program instructions executable by the at least one hardware processor to:

receive, as input, a dataset comprising a plurality of data instances,

extract a feature vector representation of each of said data instances in said dataset,

choose, from said dataset, a first data instance for adding to a subset of said dataset, wherein said first data instance is removed from said dataset,

perform an iterative process comprising:

(i) identifying one of said data instances in said dataset which represents a maximal information addition to said subset, based, at least in part, on measuring an information difference parameter between said feature vector representation of said identified data instance and said feature vector representations of all of said data instances in said subset, and

(ii) adding said identified data instance to said subset and removing said identified data instance from said dataset,

until said information difference parameter is lower than a predetermined threshold,

assign a label to each of said data instances in said subset,

train a machine learning model on a training set comprising (a) said data instances in said subset, and (b) said assigned labels, and

apply said trained machine learning model to a second subset selected from said dataset, to annotate each of said data instances in said second subset with one of said labels.

2. The system of claim 1 , wherein said program instructions are further executable to receive a budgetary constraint, wherein said budgetary constraint causes said iterative process to stop before said information difference parameter is lower than said predetermined threshold.

3. The system of claim 2 , wherein said budgetary constraint is expressed as at least one of: a maximal number of data instances in said subset, and a maximal computational time limit applicable to the performance of said repeating.

4. The system of claim 1 , wherein said feature vector representation is obtained using at least one of: supervised or unsupervised dictionary learning, autoencoding, k-means clustering, principal component analysis, independent component analysis, linear embedding, neural network model features, frequency space mapping, image histogram, and scale-invariant feature transform (SIFT) feature detection algorithm.

5. The system of claim 1 , wherein said measuring of said information difference is based, at least in part, on one of: a Euclidean distance calculation, an information entropy calculation, and a feature probability distribution calculation.

6. A computer-implemented method comprising:

receiving, as input, a dataset comprising a plurality of data instances;

extracting a feature vector representation of each of said data instances in said dataset;

choosing, from said dataset, a first data instance for adding to a subset of said dataset, wherein said first data instance is removed from said dataset;

performing an iterative process comprising:

identifying one of said data instances in said dataset which represents a maximal information addition to said subset, based, at least in part, on measuring an information difference parameter between said feature vector representation of said identified data instance and said feature vector representations of all of said data instances in said subset, and

(ii) adding said identified data instance to said subset and removing said identified data instance from said dataset,

until said information difference parameter is lower than a predetermined threshold;

assigning a label to each of said data instances in said subset,

training a machine learning model on a training set comprising (a) said data instances in said subset, and (b) said assigned labels, and

applying said trained machine learning model to a second subset selected from said dataset, to annotate each of said data instances in said second subset with one of said labels.

7. The computer-implemented method of claim 6 , further comprising receiving a budgetary constraint, wherein said budgetary constraint causes said iterative process to stop before said information difference parameter is lower than said predetermined threshold.

8. The computer-implemented method of claim 7 , wherein said budgetary constraint is expressed as at least one of: a maximal number of data instances in said subset, and a maximal computational time limit applicable to the performance of said repeating.

9. The computer-implemented method of claim 6 , wherein said feature vector representation is obtained using at least one of: supervised or unsupervised dictionary learning, autoencoding, k-means clustering, principal component analysis, independent component analysis, linear embedding, neural network model features, frequency space mapping, image histogram, and scale-invariant feature transform (SIFT) feature detection algorithm.

10. The computer-implemented method of claim 6 , wherein said measuring of said information difference is based, at least in part, on one of: a Euclidean distance calculation, an information entropy calculation, and a feature probability distribution calculation.

11. A computer program product comprising a non-transitory computer-readable storage medium having program instructions embodied therewith, the program instructions executable by at least one hardware processor to:

receive, as input, a dataset comprising a plurality of data instances;

extract a feature vector representation of each of said data instances in said dataset;

choose, from said dataset, a first data instance for adding to a subset of said dataset, wherein said first data instance is removed from said dataset;

perform an iterative process comprising:

(i) identifying one of said data instances in said dataset which represents a maximal information addition to said subset, based, at least in part, on measuring an information difference parameter between said feature vector representation of said identified data instance and said feature vector representations of all of said data instances in said subset, and

(ii) adding said identified data instance to said subset and removing said identified data instance from said dataset,

until said information difference parameter is lower than a predetermined threshold;

assign a label to each of said data instances in said subset;

train a machine learning model on a training set comprising (a) said data instances in said subset, and (b) said assigned labels; and

apply said trained machine learning model to a second subset selected from said dataset, to annotate each of said data instances in said second subset with one of said labels.

12. The computer program product of claim 11 , wherein said program instructions are further executable to receive a budgetary constraint, wherein said budgetary constraint causes said iterative process to stop before said information difference parameter is lower than said predetermined threshold.

13. The computer program product of claim 12 , wherein said budgetary constraint is expressed as at least one of: a maximal number of data instances in said subset, and a maximal computational time limit applicable to the performance of said repeating.

14. The computer program product of claim 11 , wherein said feature vector representation is obtained using at least one of: supervised or unsupervised dictionary learning, autoencoding, k-means clustering, principal component analysis, independent component analysis, linear embedding, neural network model features, frequency space mapping, image histogram, and scale-invariant feature transform (SIFT) feature detection algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2021
From: SHABTAY, OR; SHLOMO, ERAN
To: DATALOOP LTD.
Reel/Frame 056113/0846 →
Continuity (2)
Provisional Application 63019353 · May 3, 2020
Related Publication 20210342642A1 · Nov 4, 2021
Cited By (2)
US 12,524,501 US 12,554,796