IP Library Granted Patent US 11,238,365
Granted Patent B2
US 11,238,365 · App. 15/858,001 · Granted Feb 1, 2022

Method and system for detecting anomalies in data labels

Inventors: Francis Hsu (Santa Clara, CA); Mridul Jain (Sunnyvale, CA); Saurabh Tewari (Sunnyvale, CA)
Assignee: VERIZON MEDIA INC.
G06N20/00G06F16/285G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,365
App. No.
15/858,001
Granted
Feb 1, 2022
Kind
B2
Abstract

The present teaching relates to a method and system for validating labels of training data. A first group of data records associated with the training data are received, wherein each of the first group of data records includes a vector having at least one feature and a first label. For each of the first group of data records, a second label is determined based on the at least one feature in accordance with a first model. Thereafter, a loss based on the first label associated with the data record and the second label is obtained, and the data record having an incorrect first label is classified when the loss meets a pre-determined criterion. Upon classifying the data records, a sub-group of the first group of data records is generated, wherein each of the data records included in the sub-group has the incorrect first label.

Claims (61)

1. A method, implemented on a machine having at least one processor, storage, and a communication platform capable of connecting to a network for validating labels of training data, the method comprising:

receiving a first group of data records associated with the training data, wherein each of the first group of data records includes a vector having at least one feature and a first label determined via a heuristic model;

for each of the first group of data records,

determining a second label based on the at least one feature in accordance with a first model different from the heuristic model and obtained via training based on the data records;

in the event that the first label does not match the second label, obtaining a loss indicating a degree of mismatch between the first label and the second label; and

classifying the data record as having an incorrect first label when the loss meets a pre-determined criterion; and

generating a sub-group of the first group of data records, each of which has the incorrect first label.

2. The method of claim 1 , wherein the loss is determined based on a discrepancy between the first label and the second label.

3. The method of claim 1 , further comprising:

correcting, for each of the sub-group of data records, the associated first label based on the corresponding second label.

4. The method of claim 1 , further comprising:

clustering the sub-group of data records into one or more clusters; and

identifying, with respect to each of the one or more clusters, a cause that leads to the incorrect first labels.

5. The method of claim 4 , further comprising:

determining, for at least some cause identified, a correction directed to a labeling model used in generating a corresponding first label.

6. The method of claim 5 , wherein the labeling model corresponds to a heuristic labeling model associated with a threshold; and

the correction directed to the heuristic labeling model is to adjust the threshold of the heuristic labeling model.

7. The method of claim 1 , wherein the first model is derived by:

obtaining a second group of data records from the training data, wherein the first and second groups of data records are disjoint;

performing supervised learning of the first model using the second group of data records based on the first labels of the data records in the second group; and

deriving the first model when the supervised learning converges.

8. A system for validating labels of training data, comprising:

at least one processor configured to receive a first group of data records associated with the training data, wherein each of the first group of data records includes a vector having at least one feature and a first label determined via a heuristic model;

for each of the first group of data records,

determine a second label based on the at least one feature in accordance with a first model different from the heuristic model and obtained via training based on the data records;

in the event that the first label does not match the second label, obtain a loss indicating a degree of mismatch between the first label and the second label; and

classify the data record as having an incorrect first label when the loss meets a pre-determined criterion; and

generate a sub-group of the first group of data records, each of which has the incorrect first label.

9. The system of claim 8 , wherein the loss is determined based on a discrepancy between the first label and the second label.

10. The system of claim 8 , wherein the at least one processor is further configured to:

correct, for each of the sub-group of data records, the associated first label based on the corresponding second label.

11. The system of claim 8 , wherein the at least one processor is further configured to:

cluster the sub-group of data records into one or more clusters; and

identify, with respect to each of the one or more clusters, a cause that leads to the incorrect first labels.

12. The system of claim 11 , wherein the at least one processor is further configured to:

determine, for at least some cause identified, a correction directed to a labeling model used in generating a corresponding first label.

13. The system of claim 12 , wherein

the labeling model corresponds to a heuristic labeling model associated with a threshold; and

the correction directed to the heuristic labeling model is to adjust the threshold of the heuristic labeling model.

14. The system of claim 8 , wherein the at least one processor is further configured to:

obtain a second group of data records from the training data, wherein the first and second groups of data records are disjoint;

perform supervised learning of the first model using the second group of data records based on the first labels of the data records in the second group; and

derive the first model when the supervised learning converges.

15. A non-transitory computer readable medium including computer executable instructions, wherein the instructions, when executed by a computer, cause the computer to perform a method for validating labels of training data, the method comprising:

receiving a first group of data records associated with the training data, wherein each of the first group of data records includes a vector having at least one feature and a first label determined via a heuristic model;

for each of the first group of data records,

determining a second label based on the at least one feature in accordance with a first model different from the heuristic model and obtained via training based on the data records;

in the event that the first label does not match the second label, obtaining a loss indicating a degree of mismatch between the first label and the second label; and

classifying the data record as having an incorrect first label when the loss meets a pre-determined criterion; and

generating a sub-group of the first group of data records, each of which has the incorrect first label.

16. The non-transitory computer readable medium of claim 15 , wherein the loss is determined based on a discrepancy between the first label and the second label.

17. The non-transitory computer readable medium of claim 15 , the method further comprising:

correcting, for each of the sub-group of data records, the associated first label based on the corresponding second label.

18. The non-transitory computer readable medium of claim 15 , the method further comprising:

clustering the sub-group of data records into one or more clusters; and

identifying, with respect to each of the one or more clusters, a cause that leads to the incorrect first labels.

19. The non-transitory computer readable medium of claim 18 , the method further comprising:

determining, for at least some cause identified, a correction directed to a labeling model used in generating a corresponding first label.

20. The non-transitory computer readable medium of claim 19 , wherein

the labeling model corresponds to a heuristic labeling model associated with a threshold; and

the correction directed to the heuristic labeling model is to adjust the threshold of the heuristic labeling model.

Assignments (6)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
CHANGE OF NAME Recorded Feb 24, 2020
From: OATH (AMERICAS) INC.
To: VERIZON MEDIA INC.
Reel/Frame 051999/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2018
From: YAHOO HOLDINGS, INC.
To: OATH INC.
Reel/Frame 045240/0310 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2017
From: HSU, FRANCIS; JAIN, MRIDUL; TEWARI, SAURABH
To: OATH INC.
Reel/Frame 044505/0203 →
Continuity (1)
Related Publication 20190205794A1 · Jul 4, 2019