IP Library › Granted Patent US 12,572,622
Granted Patent B2
US 12,572,622 · App. 17/961,072 · Granted Mar 10, 2026

Identifying incorrect labels and improving label correction for machine learning (ML) for security

Inventors: Miao Zhang (Palo Alto, CA); Loc Bui (San Jose, CA); Dianhuan Lin (Sunnyvale, CA); Rex Shang (Los Altos, CA); Howie Xu (Palo Alto, CA)
Assignee: Zscaler, Inc.
G06F18/217G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,622
App. No.
17/961,072
Filed
Oct 6, 2022
Granted
Mar 10, 2026
Kind
B2
Art Unit
2444
USPC
713/154
Abstract

Systems and methods for identifying incorrect labels and improving label correction for machine learning for security. The systems and methods including receiving data with labels; training one or more Machine Learning (ML) models to label the received data; identifying disagreements between the labels provided by the one or more ML models and the labels received with the data; and providing one or more groups of the data for review for incorrect labels.

Claims (39)

1 . A method for identifying incorrect labels in datasets, the method comprising steps of:

receiving data with labels;

training one or more Machine Learning (ML) models to label the received data;

identifying disagreements between the labels provided by the one or more ML models and the labels received with the data; and

providing one or more groups of the data for review for incorrect labels, wherein the providing includes the steps of:

(i) assigning a priority for review to the data based on whether (a) a plurality of ML models agree with each other but disagree with the received labels, or (b) the plurality of ML models disagree with each other and also disagree with the received labels; and

(ii) further providing one or more sub-groups of the groups of data for review, wherein the sub-groups are determined based on at least one of URL domain and similarity of the data, and wherein different ML models are used to provide the groups and the sub-groups, including ML models trained using supervised learning and active learning to provide the groups of the data for review, and ML models trained using unsupervised learning to provide the sub-groups of the data for review.

2 . The method of claim 1 , wherein the providing the one or more groups of the data for review comprises assigning a priority for review to the data based on whether (i) a plurality of ML models agree with each other but disagree with the received labels, or (ii) the plurality of ML models disagree with each other and also disagree with the received labels.

3 . The method of claim 1 , wherein a plurality of ML models disagree about labels for one or more files in the received data, and wherein the one or more groups of data provided for review include the one or more files.

4 . The method of claim 1 , wherein one or more sub-groups are provided for review, and wherein the one or more sub-groups include files from the one or more groups of the data.

5 . The method of claim 4 , wherein the groups and sub-groups are provided based on one of URL domain and similarity.

6 . The method of claim 4 , wherein different ML models are used to provide the groups and the sub-groups.

7 . The method of claim 6 , wherein ML models which provide the groups of the data are trained using supervised learning and then refined using active learning to identify a set of examples to be re-labelled, while ML models which provide the sub-groups are trained using unsupervised learning.

8 . A cloud-based system for identifying incorrect labels in datasets, the cloud based system comprising:

one or more processors and memory storing instructions that, when executed, cause the one or more processors to:

receive data with labels;

train one or more Machine Learning (ML) models to label the received data;

identify disagreements between the labels provided by the one or more ML models and the labels received with the data; and

provide one or more groups of the data for review for incorrect labels, wherein the one or more groups are provided based on:

(i) assigning a priority for review to the data based on whether (a) a plurality of ML models agree with each other but disagree with the received labels, or (b) the plurality of ML models disagree with each other and also disagree with the received labels; and

(ii) further providing one or more sub-groups of the groups of data for review, wherein the sub-groups are determined based on at least one of URL domain and similarity of the data, and wherein different ML models are used to provide the groups and the sub-groups, including ML models trained using supervised learning and active learning to provide the groups of the data for review, and ML models trained using unsupervised learning to provide the sub-groups of the data for review.

9 . The cloud-based system of claim 8 , wherein the one or more groups of data are provided by assigning a priority for review to the data based on whether (i) a plurality of ML models agree with each other but disagree with the received labels, or (ii) the plurality of ML models disagree with each other and also disagree with the received labels.

10 . The method of claim 8 , wherein a plurality of ML models disagree about labels for one or more files in the received data, and wherein the one or more groups of data provided for review include the one or more files.

11 . The method of claim 8 , wherein one or more sub-groups are provided for review, and wherein the one or more sub-groups include files from the one or more groups of the data.

12 . The method of claim 11 , wherein the groups and sub-groups are provided based on one of URL domain and similarity.

13 . The method of claim 11 , wherein different ML models are used to provide the groups and the sub-groups.

14 . The method of claim 13 , wherein ML models which provide the groups of the data are trained using supervised learning then refined using active learning to identify a set of examples to be re-labelled, while ML models which provide the sub-groups are trained using unsupervised learning.

15 . A non-transitory computer-readable medium comprising instructions for identifying incorrect labels in datasets that, when executed, cause one or more processors to perform steps of:

receiving data with labels;

training one or more Machine Learning (ML) models to label the received data;

identifying disagreements between the labels provided by the one or more ML models and the labels received with the data; and

providing one or more groups of the data for review for incorrect labels, wherein the providing includes the steps of:

(i) assigning a priority for review to the data based on whether (a) a plurality of ML models agree with each other but disagree with the received labels, or (b) the plurality of ML models disagree with each other and also disagree with the received labels; and

(ii) further providing one or more sub-groups of the groups of data for review, wherein the sub-groups are determined based on at least one of URL domain and similarity of the data, and wherein different ML models are used to provide the groups and the sub-groups, including ML models trained using supervised learning and active learning to provide the groups of the data for review, and ML models trained using unsupervised learning to provide the sub-groups of the data for review.

16 . The non-transitory computer-readable medium of claim 15 , wherein the providing the one or more groups of the data for review comprises assigning a priority for review to the data based on whether (i) a plurality of ML models agree with each other but disagree with the received labels, or (ii) the plurality of ML models disagree w ac her and also disagree with the received labels.

17 . The non-transitory computer-readable medium of claim 15 , wherein a plurality of ML models disagree about labels for one or more files in the received data, and wherein the one or more groups of data provided for review include the one or more files.

18 . The non-transitory computer-readable medium of claim 15 , wherein one or more sub-groups are provided for review, and wherein the one or more sub-groups include files from the one or more groups of the data.

19 . The non-transitory computer-readable medium of claim 18 , wherein different ML models are used to provide the groups and the sub-groups.

20 . The non-transitory computer-readable medium of claim 19 , wherein ML models which provide the groups of the data are trained using supervised learning and then refined using active learning to identify a set of examples to be re-labelled, while ML models which provide the sub-groups are trained using unsupervised learning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2022
From: ZHANG, MIAO; BUI, LOC; LIN, DIANHUAN; SHANG, REX; XU, HOWIE
To: ZSCALER, INC.
Reel/Frame 061335/0939 →
Continuity (1)
Related Publication 20240119120A1 · Apr 11, 2024
References Cited (48)
US 6009475A · Shrader · 1999 [cited by applicant]
US 6138162A · Pistriotto et al. · 2000 [cited by applicant]
US 7316029B1 · Parker et al. · 2008 [cited by applicant]
US 7383569B1 · Elgressy et al. · 2008 [cited by applicant]
US 7620985B1 · Bush et al. · 2009 [cited by applicant]
US 8166533B2 · Yuan · 2012 [cited by applicant]
US 8499348B1 · Rubin · 2013 [cited by applicant]
US 8677471B2 · Karels et al. · 2014 [cited by applicant]
US 9065850B1 · Sobrier · 2015 [cited by applicant]
US 9152789B2 · Natarajan et al. · 2015 [cited by applicant]
US 9773107B2 · White et al. · 2017 [cited by applicant]
US 10142362B2 · Weith et al. · 2018 [cited by applicant]
US 10154067B2 · Smith et al. · 2018 [cited by applicant]
US 10348599B2 · O'Neil et al. · 2019 [cited by applicant]
US 10362048B2 · Alexander et al. · 2019 [cited by applicant]
US 10419477B2 · Desai et al. · 2019 [cited by applicant]
US 10439985B2 · O'Neil · 2019 [cited by applicant]
US 10498605B2 · Weith et al. · 2019 [cited by applicant]
US 10505899B1 · Singh et al. · 2019 [cited by applicant]
US 20050193222A1 · Greene · 2005 [cited by applicant]
US 20050210065A1 · Nigam · 2005 [cited by examiner]
US 20060095970A1 · Rajagopal et al. · 2006 [cited by applicant]
US 20070233477A1 · Halowani et al. · 2007 [cited by applicant]
US 20100115621A1 · Staniford et al. · 2010 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170063886A1 · Muddu et al. · 2017 [cited by applicant]
US 20170078329A1 · Hwang et al. · 2017 [cited by applicant]
US 20170272465A1 · Steele · 2017 [cited by applicant]
US 20180041471A1 · Sudo et al. · 2018 [cited by applicant]
US 20180150758A1 · Niininen et al. · 2018 [cited by applicant]
US 20180293381A1 · Tseng et al. · 2018 [cited by applicant]
US 20190281073A1 · Weith et al. · 2019 [cited by applicant]
US 20190319972A1 · Desai · 2019 [cited by applicant]
US 20190349283A1 · O'Neil et al. · 2019 [cited by applicant]
US 20200021618A1 · Smith et al. · 2020 [cited by applicant]
US 20210042645A1 · Sharma · 2021 [cited by examiner]
US 20210125106A1 · Okamoto · 2021 [cited by examiner]
US 20220121987A1 · Grady · 2022 [cited by examiner]
US 20220391719A1 · Mansour · 2022 [cited by examiner]
US 20230251856A1 · Ni · 2023 [cited by examiner]
US 20230259883A1 · Misler · 2023 [cited by examiner]
US 20240282459A1 · Pilitsis · 2024 [cited by examiner]
WO 2018152303A1 · 2018 [cited by applicant]
Jordaney, Roberto, et al., “Transcend: Detecting concept drift in malware classification models,” 26th {USENIX} Security Symposium ({USENIX} Security 17), 2017. [cited by applicant]
Kantchelian, Alex, J. D. Tygar, and Anthony Joseph, “Evasion and hardening of tree ensemble classifiers,” International Conference on Machine Learning, 2016. [cited by applicant]
Tolomei, Gabriele, et al., “Interpretable predictions of tree-based ensembles via actionable feature tweaking,” Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 20… [cited by applicant]
Aug. 13, 2019, International Preliminary Report on Patentability and Written Opinion for International Application No. PCT/US2018/015902. [cited by applicant]
Aug. 20, 2019, International Preliminary Report on Patentability and Written Opinion for International Application No. PCT/US2018/018325. [cited by applicant]