IP Library › Granted Patent US 11,580,450
Granted Patent B2
US 11,580,450 · App. 16/744,890 · Granted Feb 14, 2023

System and method for efficiently managing large datasets for training an AI model

Inventors: Robert R. Price (Palo Alto, CA); Matthew A. Shreve (Mountain View, CA)
Assignee: Palo Alto Research Center Incorporated
G06N20/00G06K9/6218G06K9/6256
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,450
App. No.
16/744,890
Filed
Jan 16, 2020
Granted
Feb 14, 2023
Kind
B2
Art Unit
2663
USPC
706/12
Abstract

Embodiments described herein provide a system for facilitating efficient dataset management. During operation, the system obtains a first dataset comprising a plurality of elements. The system then determines a set of categories for a respective element of the plurality of elements by applying a plurality of AI models to the first dataset. A respective category can correspond to an AI model. Subsequently, the system selects a set of sample elements associated with a respective category of a respective AI model and determines a second dataset based on the selected sample elements.

Claims (36)

1. A method for facilitating efficient dataset management, comprising:

obtaining a first dataset comprising a plurality of data elements for training a first artificial intelligence (AI) model;

determining respective sets of categories of data elements by applying a plurality of AI models to the first dataset, wherein a respective set of categories corresponds to an AI model;

determining a set of joint categories from the sets of categories, wherein a respective joint category corresponds to at least two categories from at least two sets of categories, respectively;

selecting a set of sample data elements associated with a respective category of a respective set of categories by obtaining the set of sample data elements from the set of joint categories;

determining a second dataset based on the selected sample data elements; and

training the first AI model using the second dataset.

2. The method of claim 1 , wherein the plurality of AI models includes one or more pre-trained classifiers.

3. The method of claim 1 , wherein applying an AI model of the plurality of AI models to the first dataset comprises categorizing the plurality of data elements into a corresponding set of categories supported by the AI model.

4. The method of claim 1 , wherein applying an AI model of the plurality of AI models to the first dataset comprises:

obtaining embeddings for the plurality of data elements based on the AI model; and

grouping the plurality of data elements into a set of clusters based on the embeddings.

5. The method of claim 4 , wherein grouping the plurality of data elements comprises applying a k-means clustering algorithm to the embeddings.

6. The method of claim 1 , further comprising determining a number of sample data elements to be selected for a respective category of a set of categories associated with a respective AI model.

7. The method of claim 6 , further comprising determining the number of sample data elements based on the set of joint categories obtained from the sets of categories.

8. The method of claim 6 , further comprising determining the number of sample data elements for a category of a set of categories associated with an AI model without considering a category of another AI model.

9. The method of claim 6 , wherein the number of sample data elements selected for the category corresponds to a proportion of data elements for the category in the first dataset.

10. The method of claim 1 , wherein training the first AI model further comprises training the first AI model using the second dataset based on proportions of elements in the first dataset in a respective category of a set of categories associated with a respective AI model.

11. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for facilitating efficient dataset management, the method comprising:

obtaining a first dataset comprising a plurality of data elements for training a first artificial intelligence (AI) model;

determining respective sets of categories of data elements by applying a plurality of AI models to the first dataset, wherein a respective set of categories corresponds to an AI model;

determining a set of joint categories from the sets of categories, wherein a respective joint category corresponds to at least two categories from at least two sets of categories, respectively;

selecting a set of sample data elements associated with a respective category of a respective set of categories by obtaining the set of sample data elements from the set of joint categories;

determining a second dataset based on the selected sample data elements; and

training the first AI model using the second dataset.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the plurality of AI models includes one or more pre-trained classifiers.

13. The non-transitory computer-readable storage medium of claim 11 , wherein applying an AI model of the plurality of AI models to the first dataset comprises categorizing the plurality of data elements into a corresponding set of categories supported by the AI model.

14. The non-transitory computer-readable storage medium of claim 11 , wherein applying an AI model of the plurality of AI models to the first dataset comprises:

obtaining embeddings for the plurality of data elements based on the AI model; and

grouping the plurality of data elements into a set of clusters based on the embeddings.

15. The non-transitory computer-readable storage medium of claim 14 , wherein grouping the plurality of data elements comprises applying a k-means clustering algorithm to the embeddings.

16. The non-transitory computer-readable storage medium of claim 11 , wherein the method further comprises determining a number of sample data elements to be selected for a respective category of a set of categories associated with a respective AI model.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the method further comprises determining the number of sample data elements based on the set of joint categories obtained from the sets of categories.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the method further comprises determining the number of sample data elements for a category of a set of categories associated with an AI model without considering a category of another AI model.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the number of sample data elements selected for the category corresponds to a proportion of data elements for the category in the first dataset.

20. The non-transitory computer-readable storage medium of claim 11 , wherein training the first AI model further comprises training the first AI model using the second dataset based on proportions of elements in the first dataset in a respective category of a set of categories associated with a respective AI model.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2026
From: XEROX CORPORATION
To: GENESEE VALLEY INNOVATIONS, LLC
Reel/Frame 075020/0755 →
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2020
From: PRICE, ROBERT R.; SHREVE, MATTHEW A.
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 051540/0756 →
Continuity (1)
Related Publication 20210224683A1 · Jul 22, 2021