IP Library Granted Patent US 11,947,633
Granted Patent B2
US 11,947,633 · App. 17/107,574 · Granted Apr 2, 2024

Oversampling for imbalanced test data

Inventors: Hongwei Shang (Sunnyvale, CA); Jean-Marc Langlois (Menlo Park, CA); Kostas Tsioutsiouliklis (Saratoga, CA); Changsung Kang (San Jose, CA)
Assignee: Yahoo Assets LLC
G06F18/24137G06F11/076G06F18/214G06F18/24155
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,947,633
App. No.
17/107,574
Granted
Apr 2, 2024
Kind
B2
Abstract

One or more computing devices, systems, and/or methods for oversampling for imbalanced test data are provided. A classifier is configured to classify data points as either belonging to a first class or a second class. A determination may be made that the first class and the second class are imbalanced where a first number of data points estimated to be part of the first class is a threshold amount less than a second number of data points estimated to be part of the second class. An oversampling ratio is determined for the first class. The oversampling ratio is used to select a sample set of data points for editorial labeling, where the sampling set of data points comprises a total number of data points below a threshold amount.

Claims (41)

1. A method, comprising:

executing, on a processor of a computing device, instructions that cause the computing device to perform operations, the operations comprising:

configuring a classifier to classify data points as either belonging to a first class or a second class;

in response to determining that the first class and the second class are imbalanced where a first number of data points estimated to be part of the first class is a threshold amount less than a second number of data points estimated to be part of the second class, determining an oversampling ratio for the first class; and

selecting, based upon the oversampling ratio, a sample set of data points for editorial labeling comprising a selected number of data points estimated to be part of the first class and a selected number of data points estimated to be part of the second class, wherein the sample set of data points comprises a total number of data points below a threshold amount.

2. The method of claim 1 , comprising:

determining the oversampling ratio based upon an imbalance ratio between the first class and the second class.

3. The method of claim 2 , comprising:

determining the imbalance ratio based upon a ratio of the first number of data points and the second number of data points.

4. The method of claim 1 , comprising:

determining the oversampling ratio based a production specification.

5. The method of claim 4 , wherein the production specification corresponds to an error margin for a precision metric.

6. The method of claim 4 , wherein the production specification corresponds to an error margin for a recall metric.

7. The method of claim 1 , comprising

determining a precision confidence interval corresponding to a confidence that a precision metric for the classifier is correct.

8. The method of claim 7 , comprising:

in response to determining that a confidence of the precision confidence interval being correct is below a threshold, generating a precision distribution.

9. The method of claim 1 , comprising

determining a recall confidence interval corresponding to a confidence that a recall metric for the classifier is correct.

10. The method of claim 9 , comprising:

in response to determining that a confidence of the recall confidence interval being correct is below a threshold, generating a recall distribution.

11. The method of claim 1 , comprising:

generating an approximate distribution of a precision metric based upon at least one of a frequentists technique or a Bayesian posterior distribution technique.

12. The method of claim 1 , comprising:

generating an approximate distribution of a recall metric based upon at least one of a frequentists technique or a Bayesian posterior distribution technique.

13. A non-transitory machine readable medium having stored thereon processor-executable instructions that when executed cause performance of operations, the operations comprising:

identifying a first class and a second class of data points that are classified by a classifier;

in response to determining that the first class and the second class are imbalanced where a first number of data points estimated to be part of the first class is a threshold amount less than a second number of data points estimated to be part of the second class, determining an oversampling ratio for the first class; and

selecting, based upon the oversampling ratio, a sample set of data points for editorial labeling comprising a selected number of data points estimated to be part of the first class and a selected number of data points estimated to be part of the second class, wherein the sample set of data points comprises a total number of data points below a threshold amount.

14. The non-transitory machine readable medium of claim 13 , wherein the operations comprise:

performing a simulation utilizing a bootstrap method corresponding to confidence intervals associated with a precision metric and a recall metric.

15. The non-transitory machine readable medium of claim 13 , wherein the operations comprise:

performing a simulation utilizing a Monte-Carlo method corresponding to confidence intervals associated with a precision metric and a recall metric.

16. A computing device comprising:

a processor; and

memory comprising processor-executable instructions that when executed by the processor cause performance of operations, the operations comprising:

determining that a first class and a second class of data points classified by a classifier are imbalanced where a first number of data points estimated to be part of the first class is a threshold amount less than a second number of data points estimated to be part of the second class,

calculating an oversampling ratio for the first class; and

selecting, based upon the oversampling ratio, a sample set of data points for editorial labeling comprising a selected number of data points estimated to be part of the first class and a selected number of data points estimated to be part of the second class, wherein the sample set of data points comprises a total number of data points below a threshold amount.

17. The computing device of claim 16 , wherein the operations comprise:

determining a prediction confidence interval for a precision metric based upon a binominal distribution of predictive positive data.

Assignments (3)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2020
From: SHANG, HONGWEI; LANGLOIS, JEAN-MARC; TSIOUTSIOULIKLIS, KOSTAS; KANG, CHANGSUNG
To: VERIZON MEDIA INC.
Reel/Frame 054495/0557 →