IP Library Granted Patent US 10,848,508
Granted Patent B2
US 10,848,508 · App. 15/985,644 · Granted Nov 24, 2020

Method and system for generating synthetic feature vectors from real, labelled feature vectors in artificial intelligence training of a big data machine to defend

Inventors: Victor Chen (San Jose, CA); Ignacio Arnaldo (San Jose, CA); Constantinos Bassias (San Jose, CA)
H04L63/1425G06F21/552G06K9/6255G06K9/6259G06N3/0454G06N3/0472G06N3/084G06N7/00G06N20/00G06K9/6247G06K9/6262
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,848,508
App. No.
15/985,644
Granted
Nov 24, 2020
Kind
B2
Abstract

Identifying and detecting threats to an enterprise system groups log lines from enterprise data sources and/or from incoming data traffic. The process applies artificial intelligence processing to the statistical outlier in the event of the statistical outliers comprises a sparsely labelled real data set, by receiving the sparsely labelled real data set for identifying malicious data and comprising real labelled feature vectors and generating a synthetic data set comprising a plurality of synthetic feature vectors derived from the real, labelled feature vectors. The process further identifies the sparsely labelled real data set as a local data set and the synthetic data set as a global set. The process further applies a transfer learning framework for mixing the global data set with the local data set for increasing the precision recall area under curve (PR AUC) for reducing false positive indications occurring in analysis of the threats to the enterprise.

Claims (58)

1. A method for identifying and detecting threats to an enterprise or e-commerce system, the method comprising:

grouping log lines belonging to one or more log line parameters from one or more enterprise or e-commerce system data sources or from incoming data traffic to the enterprise or e-commerce system;

extracting one or more features from the grouped log lines into one or more features tables;

using one or more statistical outlier detection methods on the one or more features tables to identify statistical outliers;

labeling, in response to received instructions, the statistical outliers to create one or more labeled features tables;

using the one or more labeled features tables to create an adaptive rules model for further identification of statistical outliers;

when identified labeled statistical outliers comprise a sparsely labeled real data set, applying artificial intelligence processing to said identified labeled statistical outliers, comprising the steps of:

receiving said sparsely labelled real data set for identifying malicious data and comprising real, labelled feature vectors;

generating a synthetic data set comprising a plurality of synthetic feature vectors derived from said real, labelled feature vectors;

identifying said sparsely labelled real data set as a local data set and said synthetic data set as a global set;

applying a transfer learning framework for mixing said global data set with said local data set for increasing the precision recall area under curve (PR AUC) for reducing false positive indications occurring the in analysis of said threats to the enterprise; and

preventing access by various threats to the enterprise or e-commerce system and detecting threats to the enterprise or e-commerce system in real-time based on a model formed using the generated synthetic data set.

2. The method of claim 1 , further comprising the steps of labeling the output of a single top scores vector, and said adaptive rules model to create at least one labeled features matrix for providing new input to a supervised learning module for updating one or more identified threat labels.

3. The method of claim 1 , further comprising the step of refining said adaptive rules model for identifying statistical outliers and preventing access to said enterprise system of categorized threats by detecting new threats in real time and reducing a time elapsed between threat detection of the enterprise system.

4. The method of claim 1 , further comprising the step of generating negative labels for training and evaluation by designating unlabeled feature vectors as negative and randomly sampling unlabeled feature vectors within a predetermined date range corresponding to a date range of existing positive samples.

5. The method of claim 1 , further comprising the step of generating and using synthetic feature vectors from real, labelled feature vectors for resolving data sparsity limitations when modeling anomalous events.

6. The method of claim 1 , further comprising the step of using a transfer learning framework for simulating the effect of mixing synthetic and real feature vectors for forming a guideline in the absence of online A/B testing for evaluating experimental models for identifying malicious data on fresh, new datasets.

7. An apparatus for training a big data machine to defend an enterprise system, the apparatus comprising:

one or more processors;

system memory coupled to the one or more processors;

one or more non-transitory memory units coupled to the one or more processors; and

threat identification and detection code stored on the one or more non-transitory memory units that when executed by the one or more processors are configured to perform a method, the method comprising:

grouping log lines belonging to one or more log line parameters from one or more enterprise or e-commerce system data sources or from incoming data traffic to the enterprise or e-commerce system;

extracting one or more features from the grouped log lines into one or more features tables;

using the one or more labeled features tables to create an adaptive rules model for identification of statistical outliers; labeling, in response to received instructions, the statistical outliers to create one or more labeled features tables; using the one or more labeled features tables to create one or more rules for further modifying the adaptive rules model for identification of statistical outliers;

when identified labeled statistical outliers comprise a sparsely labeled real data set, applying artificial intelligence processing to said identified labeled statistical outliers, comprising the steps of:

receiving said sparsely labelled real data set for identifying malicious data and comprising real labelled feature vectors;

generating a synthetic data set comprising a plurality of synthetic feature vectors derived from said real, labelled feature vectors;

identifying said sparsely labelled real data set as a local data set and said synthetic data set as a global set;

applying a transfer learning framework for mixing said global data set with said local data set for increasing the precision recall area under curve (PR AUC) for reducing false positive indications occurring the in analysis of said threats to the enterprise; and

preventing access by various threats to the enterprise or e-commerce system and detecting threats to the enterprise or e-commerce system in real-time based on a model formed using the generated synthetic data set.

8. The apparatus of claim 7 , further comprising the steps of labeling the output of a single top scores vector, and said adaptive rules model to create at least one labeled features matrix for providing new input to a supervised learning module for updating one or more identified threat labels.

9. The apparatus of claim 7 , further comprising the step of refining said adaptive rules model for identifying statistical outliers and preventing access to said enterprise system of categorized threats by detecting new threats in real time and reducing a time elapsed between threat detection of the enterprise system.

10. The apparatus of claim 7 , wherein said threat identification and detection code stored on the one or more non-transitory memory units perform a method that further comprises the step of generating negative labels for training and evaluation by designating unlabeled feature vectors as negative and randomly sampling unlabeled feature vectors within a predetermined date range corresponding to a date range of existing positive samples.

11. The apparatus of claim 7 , further comprising the step of generating and using synthetic feature vectors from real, labelled feature vectors for resolving data sparsity limitations when modeling anomalous events.

12. The apparatus of claim 7 , wherein said threat identification and detection code stored on the one or more non-transitory memory units perform a method that further comprises the step of using a transfer learning framework for simulating the effect of mixing synthetic and real feature vectors for forming a guideline in the absence of online A/B testing for evaluating experimental models for identifying malicious data on fresh, new datasets.

13. An enterprise system for providing networked computing services to a large enterprise, said enterprise system, comprising:

an apparatus for training a big data machine to defend an enterprise system, said apparatus comprising:

one or more processors;

system memory coupled to the one or more processors;

one or more non-transitory memory units coupled to the one or more processors; and

threat identification and detection code stored on the one or more non-transitory memory units that when executed by the one or more processors are configured to perform a method, the method comprising:

grouping log lines belonging to one or more log line parameters from one or more enterprise or e-commerce system data sources or from incoming data traffic to the enterprise or e-commerce system;

extracting one or more features from the grouped log lines into one or more features tables;

using the one or more labeled features tables to create an adaptive rules model for further identification of statistical outliers;

labeling, in response to received instructions, the statistical outliers to create one or more labeled features tables;

using the one or more labeled features tables to create one or more rules for further modifying the adaptive rules model for identification of statistical outliers;

when identified labeled statistical outliers comprise a sparsely labeled real data set, applying artificial intelligence processing to said identified labeled statistical outliers, comprising the steps of:

receiving said sparsely labelled real data set for identifying malicious data and comprising real labelled feature vectors;

generating a synthetic data set comprising a plurality of synthetic feature vectors derived from said real, labelled feature vectors;

identifying said sparsely labelled real data set as a local data set and said synthetic data set as a global set;

applying a transfer learning framework for mixing said global data set with said local data set for increasing the precision recall area under curve (PR AUC) for reducing false positive indications occurring the in analysis of said threats to the enterprise; and

preventing access by various threats to the enterprise or e-commerce system and detecting threats to the enterprise or e-commerce system in real-time based on a model formed using the generated synthetic data set.

14. The enterprise system of claim 13 , further comprising the steps of labeling the output of a single top scores vector, and said adaptive rules model to create at least one labeled features matrix for providing new input to a supervised learning module for updating one or more identified threat labels.

15. The enterprise system of claim 13 , further comprising the step of refining said adaptive rules model for identifying statistical outliers and preventing access to said enterprise system of categorized threats by detecting new threats in real time and reducing a time elapsed between threat detection of the enterprise system.

16. The enterprise system of claim 13 , wherein said apparatus for training a big data machine further comprises threat identification and detection code stored on the one or more non-transitory memory units for performing a method that further comprises the step of generating negative labels for training and evaluation by designating unlabeled feature vectors as negative and randomly sampling unlabeled feature vectors within a predetermined date range corresponding to a date range of existing positive samples.

17. The enterprise system of claim 13 , further comprising the step of generating and using synthetic feature vectors from real, labelled feature vectors for resolving data sparsity limitations when modeling anomalous events.

18. The enterprise system of claim 13 , wherein said apparatus for training a big data machine further comprises threat identification and detection code stored on the one or more non-transitory memory units for performing a method that further comprises the step of using a transfer learning framework for simulating the effect of mixing synthetic and real feature vectors for forming a guideline in the absence of online A/B testing for evaluating experimental models for identifying malicious data on fresh, new datasets.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Aug 12, 2022
From: TRIPLEPOINT CAPITAL LLC
To: PATTERNEX, INC.
Reel/Frame 060797/0217 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2022
From: PATTERNEX, INC.
To: CORELIGHT, INC.
Reel/Frame 060798/0647 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2020
From: ARNALDO, IGNACIO; ARUN, ANKIT; LAM, MEI; BASSIAS, COSTAS
To: PATTERNEX, INC.
Reel/Frame 053326/0461 →
SECURITY INTEREST Recorded Nov 19, 2018
From: PATTERNEX, INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 047536/0116 →
Continuity (9)
Continuation In Part 15821231 · Nov 22, 2017
Continuation In Part 15662323 · Jul 28, 2017
Continuation In Part 15985644
Continuation In Part 15612388 · Jun 2, 2017
Continuation In Part 15985644
Continuation In Part 15382413 · Dec 16, 2016
Continuation In Part 15985644
Continuation In Part 15258797 · Sep 7, 2016
Related Publication 20190132343A1 · May 2, 2019
Cited By (5)
US 12,190,525 US 12,417,624 US 12,443,865 US 12,561,686 US 12,626,104