IP Library › Granted Patent US 11,362,910
Granted Patent B2
US 11,362,910 · App. 16/037,857 · Granted Jun 14, 2022

Distributed machine learning for anomaly detection

Inventors: Jian Lin (Apharetta, GA); Matthew Elsner (Dunwoody, GA); Ronald Williams (Austin, TX); Michael Josiah Bolding (Smyrna, GA); Yun Pan (Roswell, GA); Paul Sherwood Taylor (Redwood City, CA); Cheng-Ta Lee (Taipei, TW)
Assignee: International Business Machines Corporation
H04L41/28G06F17/15G06N20/00H04L41/16H04L63/104H04L63/1425H04L63/20H04L43/04H04L63/1433
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,362,910
App. No.
16/037,857
Granted
Jun 14, 2022
Kind
B2
Abstract

A tiered machine learning-based infrastructure comprises a first machine learning (ML) tier configured to execute within an enterprise network environment and that learns statistics for a set of use cases locally, and to alert deviations from the learned distributions. Use cases typically are independent from one another. A second machine learning tier executes external to the enterprise network environment and provides further learning support, e.g., by determining a correlation among multiple independent use cases that are running locally in the first tier. Preferably, the second tier executes in a cloud compute environment for scalability and performance.

Claims (38)

1. A method for anomaly detection in association with an enterprise environment, comprising:

providing first machine learning to train at least a first and a second analytic, the first analytic corresponding to a first use case, and the second analytic corresponding to a second use case distinct from the first use case;

outputting to a second machine learning, and in a continuous manner, anomaly information derived from the first machine learning, the anomaly information including data points detected by the first machine learning as a result of applying the first and second analytics with respect to the first and second use cases; and

providing the second machine learning based on the anomaly information to capture a correlation among observed parameters in at least the first and second use cases;

wherein the first machine learning takes place in an enterprise network, and the second machine learning takes place in a cloud computing environment distinct from the enterprise network.

2. The method as described in claim 1 wherein each of the first and second use cases uses a distinct training data set.

3. The method as described in claim 1 wherein training data for at least one of the first and second use cases is time-series data.

4. The method as described in claim 1 further including outputting configuration and training data from the enterprise network to the cloud computing environment.

5. The method as described in claim 1 wherein the correlation is a multi-dimensional distance measure.

6. The method as described in claim 1 further including taking an action with respect to detected network activity or user behavior based the captured correlation provided by the second machine learning.

7. An apparatus, comprising:

hardware processors;

computer memory holding computer program instructions executed by the hardware processors for anomaly detection in association with an enterprise environment, the computer program instructions configured to:

provide first machine learning to train at least a first and a second analytic, the first analytic corresponding to a first use case, and the second analytic corresponding to a second use case distinct from the first use case;

output to a second machine learning, and in a continuous manner, anomaly information derived from the first machine learning, the anomaly information including data points detected by the first machine learning as a result of applying the first and second analytics with respect to the first and second use cases; and

provide the second machine learning based on the anomaly information to capture a correlation among observed parameters in at least the first and second use cases;

wherein the first machine learning takes place in an enterprise network, and the second machine learning takes place in a cloud computing environment distinct from the enterprise network.

8. The apparatus as described in claim 7 wherein each of the first and second use cases uses a distinct training data set.

9. The apparatus as described in claim 7 wherein training data for at least one of the first and second use cases is time-series data.

10. The apparatus as described in claim 7 wherein the computer program instructions are further configured to output configuration and training data from the enterprise network to the cloud computing environment.

11. The apparatus as described in claim 7 wherein the computer program instructions are configured to execute a multi-dimensional distance measure to determine the correlation.

12. The apparatus as described in claim 7 wherein the computer program instructions are further configured to take an action with respect to detected network activity or user behavior based the captured correlation provided by the second machine learning.

13. A computer program product in a non-transitory computer readable medium for use in first and second data processing systems for anomaly detection in association with an enterprise environment, the computer program product holding computer program instructions that, when executed by a respective one of the first and second data processing systems, are configured to:

provide first machine learning to train at least a first and a second analytic, the first analytic corresponding to a first use case, and the second analytic corresponding to a second use case distinct from the first use case;

output to a second machine learning, and in a continuous manner, anomaly information derived from the first machine learning, the anomaly information including data points detected by the first machine learning as a result of applying the first and second analytics with respect to the first and second use cases; and

provide the second machine learning based on the anomaly information to capture a correlation among observed parameters in at least the first and second use cases;

wherein the first machine learning takes place in the first data processing system within an enterprise network, and the second machine learning takes place in the second data processing system located in a cloud computing environment distinct from the enterprise network.

14. The computer program product as described in claim 13 wherein each of the first and second use cases uses a distinct training data set.

15. The computer program product as described in claim 13 wherein training data for at least one of the first and second use cases is time-series data.

16. The computer program product as described in claim 13 wherein the computer program instructions are further configured to output configuration and training data from the first data processing system to the second data processing system.

17. The computer program product as described in claim 13 wherein the computer program instructions are configured to execute a multi-dimensional distance measure to determine the correlation.

18. The computer program product as described in claim 13 wherein the computer program instructions are further configured to take an action with respect to detected network activity or user behavior based the captured correlation provided by the second machine learning.

19. A machine learning system for anomaly detection, comprising:

a memory;

a first machine learning system executing in a first operating environment, the first machine learning system training at least a first and a second analytic, the first analytic corresponding to a first use case, and the second analytic corresponding to a second use case distinct from the first use case, the first machine learning system training on the first and second use cases simultaneously; and

a second machine learning system executing in a second operating environment remote from the first machine learning system, the second machine learning system configured to capture a multi-dimensional distance measure correlation among observed parameters in at least the first and second use cases, the observed parameters being derived by the first machine learning system;

wherein the observed parameters are output from the first machine learning system to the second machine learning system in a continuous manner.

20. The machine learning system as described in claim 19 wherein the first machine learning system executes as an application in a Security Event and Incident Management Platform (SIEM), and wherein the second machine learning system executes as an application in a cloud compute infrastructure with the second operating environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2018
From: LIN, JIAN; ELSNER, MATTHEW; WILLIAMS, RONALD; BOLDING, MICHAEL JOSEPH; PAN, YUN; LEE, CHENG-TA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046374/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2018
From: TAYLOR, PAUL SHERWOOD
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046376/0100 →
Continuity (1)
Related Publication 20200028862A1 · Jan 23, 2020
Cited By (4)
US 12,284,087 US 12,309,039 US 12,530,255 US 12,531,773