IP Library Granted Patent US 12,192,076
Granted Patent B2
US 12,192,076 · App. 17/567,310 · Granted Jan 7, 2025

Network traffic identification using machine learning

Inventors: Alexander Frazier (San Jose, CA); Dianhuan Lin (Sunnyvale, CA); Amir Levy (San Jose, CA); Amanda Carter (Dyrham, GB); Piyush Gour (Bangalore, IN)
Assignee: Zscaler, Inc.
H04L43/04H04L41/16H04L43/062H04L47/2483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,076
App. No.
17/567,310
Granted
Jan 7, 2025
Kind
B2
Abstract

Systems and methods include obtaining historical data of traffic for a plurality of locations for a cloud service; labeling the historical data as one of human and server based on a plurality of features; and utilizing the labeled historical data to train a machine learning model to classify traffic as one of human and server. The steps can further include utilizing the trained machine learning model to classify unauthenticated traffic, for the cloud service, from a specific location or a specific IP address.

Claims (36)

1. A non-transitory computer-readable storage medium having computer-readable code stored thereon for programming one or more processors to perform steps of:

obtaining historical data of traffic for a plurality of locations for a cloud service;

labeling the historical data as one of human and server based on a plurality of features;

utilizing the labeled historical data to train a machine learning model to classify traffic as one of human and server;

monitoring traffic and labeling the monitored traffic as originating from any of authenticated users and unauthenticated locations, wherein the traffic is labeled as originating from an unauthenticated location based on the traffic being tunneled via a tunnel originating from a location, the tunnel including all traffic from the location and not identifying each user or server within the location; and

utilizing the trained machine learning model to classify traffic originating from the unauthenticated locations, wherein the traffic originating from the unauthenticated locations is given a server score.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the steps include

utilizing the trained machine learning model to assign a server score to a specific location of the plurality of locations; and

responsive to the specific location being assigned an intermediate server score from the trained machine learning model, utilizing the machine learning model to perform a more granular analysis comprising analyzing individual Internet Protocol (IP) addresses associated with the specific location.

3. The non-transitory computer-readable storage medium of claim 2 , wherein the trained machine learning model classifies the traffic labeled as originating from unauthenticated locations as a split between human and server for an entire location.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the plurality of features include social networking traffic being labeled as human.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the plurality of features include daily activity with activity around a day being labeled as server and activity within business hours being labeled as human.

6. The non-transitory computer-readable storage medium of claim 1 , wherein the plurality of features include number of days active with activity every day being labeled as server and activity only during business days being labeled as human.

7. The non-transitory computer-readable storage medium of claim 1 , wherein the plurality of features include number of unique hostnames visited by the traffic where less unique hostnames are labeled as server and more unique hostnames are labeled as human.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the plurality of features include distinct applications visited by the traffic where traffic visiting less distinct applications are labeled as server and traffic visiting more distinct applications are labeled as human.

9. The non-transitory computer-readable storage medium of claim 1 , wherein the machine learning model utilizes Gradient-boosted decision trees.

10. The non-transitory computer-readable storage medium of claim 1 , wherein the steps include

utilizing the trained machine learning model to assign a server score to one or more specific IP addresses associated with a location of the plurality of locations.

11. A method comprising the steps of:

obtaining historical data of traffic for a plurality of locations for a cloud service;

labeling the historical data as one of human and server based on a plurality of features;

utilizing the labeled historical data to train a machine learning model to classify traffic as one of human and server;

monitoring traffic and labeling the monitored traffic as originating from any of authenticated users and unauthenticated locations, wherein the traffic is labeled as originating from an unauthenticated location based on the traffic being tunneled via a tunnel originating from a location, the tunnel including all traffic from the location and not identifying each user or server within the location; and

utilizing the trained machine learning model to classify traffic originating from the unauthenticated locations, wherein the traffic originating from the unauthenticated locations is given a server score.

12. The method of claim 11 , wherein the steps include

utilizing the trained machine learning model to assign a server score to a specific location of the plurality of locations; and

responsive to the specific location being assigned an intermediate server score from the trained machine learning model, utilizing the machine learning model to perform a more granular analysis comprising analyzing individual Internet Protocol (IP) addresses associated with the specific location.

13. The method of claim 12 , wherein the trained machine learning model classifies the traffic labeled as originating from unauthenticated locations as a split between human and server for an entire location.

14. The method of claim 11 , wherein the plurality of features include social networking traffic being labeled as human.

15. The method of claim 11 , wherein the plurality of features include daily activity with activity around a day being labeled as server and activity within business hours being labeled as human.

16. The method of claim 11 , wherein the plurality of features include number of days active with activity every day being labeled as server and activity only during business days being labeled as human.

17. The method of claim 11 , wherein the plurality of features include number of unique hostnames visited by the traffic where less unique hostnames are labeled as server and more unique hostnames are labeled as human.

18. The method of claim 11 , wherein the plurality of features include distinct applications visited by the traffic where traffic visiting less distinct applications are labeled as server and traffic visiting more distinct applications are labeled as human.

19. The method of claim 11 , wherein the machine learning model utilizes Gradient-boosted decision trees.

20. The method of claim 11 , wherein the steps include

utilizing the trained machine learning model to assign a server score to one or more specific IP addresses associated with a location of the plurality of locations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2022
From: FRAZIER, ALEXANDER; LIN, DIANHUAN; LEVY, AMIR; CARTER, AMANDA; GOUR, PIYUSH
To: ZSCALER, INC.
Reel/Frame 058527/0058 →
Priority Claims (1)
IN 202111053071 · Nov 18, 2021 · national
Continuity (1)
Related Publication 20230155902A1 · May 18, 2023
References Cited (46)
US 6009475A · Shrader · 1999 [cited by applicant]
US 6138162A · Pistriotto et al. · 2000 [cited by applicant]
US 7316029B1 · Parker et al. · 2008 [cited by applicant]
US 7383569B1 · Elgressy et al. · 2008 [cited by applicant]
US 7620985B1 · Bush et al. · 2009 [cited by applicant]
US 8166533B2 · Yuan · 2012 [cited by applicant]
US 8499348B1 · Rubin · 2013 [cited by applicant]
US 8677471B2 · Karels et al. · 2014 [cited by applicant]
US 9065850B1 · Sobrier · 2015 [cited by applicant]
US 9152789B2 · Natarajan et al. · 2015 [cited by applicant]
US 9407652B1 · Kesin · 2016 [cited by examiner]
US 9773107B2 · White et al. · 2017 [cited by applicant]
US 10142362B2 · Weith et al. · 2018 [cited by applicant]
US 10154067B2 · Smith et al. · 2018 [cited by applicant]
US 10348599B2 · O'Neil et al. · 2019 [cited by applicant]
US 10362048B2 · Alexander et al. · 2019 [cited by applicant]
US 10419477B2 · Desai et al. · 2019 [cited by applicant]
US 10439985B2 · O'Neil · 2019 [cited by applicant]
US 10498605B2 · Weith et al. · 2019 [cited by applicant]
US 10505899B1 · Singh et al. · 2019 [cited by applicant]
US 10836309B1 · Trundle · 2020 [cited by examiner]
US 20050193222A1 · Greene · 2005 [cited by applicant]
US 20060095970A1 · Rajagopal et al. · 2006 [cited by applicant]
US 20070233477A1 · Halowani et al. · 2007 [cited by applicant]
US 20100115621A1 · Staniford et al. · 2010 [cited by applicant]
US 20160344770A1 · Verma et al. · 2016 [cited by applicant]
US 20170063886A1 · Muddu et al. · 2017 [cited by applicant]
US 20170078329A1 · Hwang et al. · 2017 [cited by applicant]
US 20170272465A1 · Steele · 2017 [cited by applicant]
US 20170329966A1 · Koganti · 2017 [cited by examiner]
US 20170374091A1 · Igbe · 2017 [cited by examiner]
US 20180041471A1 · Sudo et al. · 2018 [cited by applicant]
US 20180150758A1 · Niininen et al. · 2018 [cited by applicant]
US 20180293381A1 · Tseng et al. · 2018 [cited by applicant]
US 20190281073A1 · Weith et al. · 2019 [cited by applicant]
US 20190319972A1 · Desai · 2019 [cited by applicant]
US 20190349283A1 · O'Neil et al. · 2019 [cited by applicant]
US 20200021618A1 · Smith et al. · 2020 [cited by applicant]
US 20200382535A1 · Herley · 2020 [cited by examiner]
US 20220129318A1 · Qiu · 2022 [cited by examiner]
WO 2018152303A1 · 2018 [cited by applicant]
Jordaney, Roberto, et al., “Transcend: Detecting concept drift in malware classification models,” 26th {USENIX} Security Symposium ({USENIX} Security 17), 2017. [cited by applicant]
Kantchelian, Alex, J. D. Tygar, and Anthony Joseph, “Evasion and hardening of tree ensemble classifiers,” International Conference on Machine Learning, 2016. [cited by applicant]
Tolomei, Gabriele, et al., “Interpretable predictions of tree-based ensembles via actionable feature tweaking,” Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 20… [cited by applicant]
Aug. 13, 2019, International Preliminary Report on Patentability and Written Opinion for International Application No. PCT/US2018/015902. [cited by applicant]
Aug. 20, 2019, International Preliminary Report on Patentability and Written Opinion for International Application No. PCT/US2018/018325. [cited by applicant]