IP Library Granted Patent US 11,714,802
Granted Patent B2
US 11,714,802 · App. 17/221,661 · Granted Aug 1, 2023

Using multiple trained models to reduce data labeling efforts

Inventors: Matthew Shreve (Campbell, CA); Francisco E. Torres (San Jose, CA); Raja Bala (Pittsford, NY); Robert R. Price (Palo Alto, CA); Pei Li (San Jose, CA)
Assignee: Palo Alto Research Center Incorporated
G06F16/2379G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,802
App. No.
17/221,661
Granted
Aug 1, 2023
Kind
B2
Abstract

A method of labeling a dataset of input samples for a machine learning task includes selecting a plurality of pre-trained machine learning models that are related to a machine learning task. The method further includes processing a plurality of input data samples through each of the pre-trained models to generate a set of embeddings. The method further includes generating a plurality of clusterings from the set of embeddings. The method further includes analyzing, by a processing device, the plurality of clusterings to extract superclusters. The method further includes assigning pseudo-labels to the input samples based on analysis.

Claims (52)

1. A method, comprising:

selecting a plurality of pre-trained machine learning models that are related to a machine learning task;

inputting a plurality of unlabeled input data samples into each of the plurality of pre-trained models to generate a set of embeddings, wherein each embedding of the set describes a feature output by the plurality of pre-trained machine learning models at multiple layer depths;

generating a plurality of clusterings from the set of embeddings;

analyzing, by a processing device, the plurality of clusterings to determine a degree of cluster member overlap between multiple of the set of embeddings;

extracting superclusters based on the degree of member overlap, wherein extracting superclusters comprises identifying a specified number of the plurality of clusterings that have a highest degree of member overlap; and

assigning pseudo-labels to the input samples based on the input samples being members of one of the superclusters.

2. The method of claim 1 , wherein the machine learning task is at least one of: an image classification task, a video classification task, or an audio classification task.

3. The method of claim 1 , further comprising generating the plurality of clusterings from a plurality of layer depths within the plurality of pre-trained models.

4. The method of claim 1 , wherein the analyzing comprises determining at least one of a structure or pattern of each of the plurality of clusterings.

5. The method of claim 1 , further comprising determining that different samples from the unlabeled input data samples belong to a same unknown class.

6. The method of claim 1 , further comprising:

receiving a human-labeled sample corresponding to a first input sample; and

labeling a second input based on the human-labeled sample and a pseudo-label associated with the first input sample and the second input sample.

7. The method of claim 1 , further comprising:

receiving a human-labeled sample corresponding to a first input sample; and

analyzing the plurality of clusterings to extract the superclusters based on the human-labeled sample.

8. A system comprising:

a memory to store a plurality of unlabeled input data samples; and

a processing device, operatively coupled to the memory, to:

select a plurality of pre-trained machine learning models that are related to a machine learning task;

input the plurality of unlabeled input data samples into each of the plurality of pre-trained models to generate a set of embeddings,

wherein each embedding of the set describes a feature output by the plurality of pre-trained machine learning models at multiple layer depths;

generate a plurality of clusterings from the set of embeddings;

analyze the plurality of clusterings to determine a degree of cluster member overlap between multiple of the set of embeddings;

extract superclusters based on the degree of member overlap, wherein extracting superclusters comprises identifying a specified number of the plurality of clusterings that have a highest degree of member overlap; and

assign pseudo-labels to the input samples based on the input samples being members of one of the superclusters.

9. The system of claim 8 , wherein the machine learning task is at least one of: an image classification task, a video classification task, or an audio classification task.

10. The system of claim 8 , the processing device further to generate the plurality of clusterings from a plurality of layer depths of a deep neural network within the plurality of pre-trained models.

11. The system of claim 8 , wherein to analyze the processing device is further to determine at least one of a structure or pattern of each of the plurality of clusterings.

12. The system of claim 8 , the processing device further to determine that different samples from the unlabeled input data samples belong to a same unknown class.

13. The system of claim 8 , the processing device further to:

receive a human-labeled sample corresponding to a first input sample; and

label a second input based on the human-labeled sample and a pseudo-label associated with the first input sample and the second input sample.

14. The system of claim 8 , the processing device further to

receive a human-labeled sample corresponding to a first input sample; and

analyze the plurality of clusterings to extract the superclusters based on the human-labeled sample.

15. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by a processing device, cause the processing device to:

select a plurality of pre-trained machine learning models that are related to a machine learning task;

input the plurality of unlabeled input data samples into each of the plurality of pre-trained models to generate a set of embeddings, wherein each embedding of the set describes a feature output by the plurality of pre-trained machine learning models at multiple layer depths;

generate a plurality of clusterings from the set of embeddings;

analyze, by the processing device, the plurality of clusterings to extract superclusters; and

assign pseudo-labels to the input samples based on analysis.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the machine learning task is at least one of: an image classification task, a video classification task, or an audio classification task.

17. The non-transitory computer-readable storage medium of claim 15 , the processing device further to generate the plurality of clusterings from a plurality of layer depths within the plurality of pre-trained models.

18. The non-transitory computer-readable storage medium of claim 15 , the processing device further to determine that different samples from the unlabeled input data samples belong to a same unknown class.

19. The non-transitory computer-readable storage medium of claim 15 , the processing device further to:

receive a human-labeled sample corresponding to a first input sample; and

label a second input based on the human-labeled sample and a pseudo-label associated with the first input sample and the second input sample.

20. The non-transitory computer-readable storage medium of claim 15 , the processing device further to

receive a human-labeled sample corresponding to a first input sample; and

analyze the plurality of clusterings to extract the superclusters based on the human-labeled sample.

Assignments (5)
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2021
From: SHREVE, MATTHEW; TORRES, FRANCISCO E.; BALA, RAJA; PRICE, ROBERT R.; LI, PEI
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 055919/0619 →
Continuity (1)
Related Publication 20220318229A1 · Oct 6, 2022
Cited By (1)
US 12,417,222