IP Library Granted Patent US 10,826,934
Granted Patent B2
US 10,826,934 · App. 15/402,503 · Granted Nov 3, 2020

Validation-based determination of computational models

Inventors: Sven Krasser (Los Angeles, CA); David Elkind (Arlington, VA); Brett Meyer (Alpharetta, GA); Patrick Crenshaw (Atlanta, GA)
Assignee: CrowdStrike, Inc.
H04L63/145G06F21/56G06N20/00H04L63/1416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,826,934
App. No.
15/402,503
Granted
Nov 3, 2020
Kind
B2
Abstract

Example techniques described herein determine a validation dataset, determine a computational model using the validation dataset, or determine a signature or classification of a data stream such as a file. The classification can indicate whether the data stream is associated with malware. A processing unit can determine signatures of individual training data streams. The processing unit can determine, based at least in part on the signatures and a predetermined difference criterion, a training set and a validation set of the training data streams. The processing unit can determine a computational model based at least in part on the training set. The processing unit can then operate the computational model based at least in part on a trial data stream to provide a trial model output. Some examples include determining the validation set based at least in part on the training set and the predetermined criterion for difference between data streams.

Claims (51)

1. A method comprising, under control of at least one processing unit:

determining respective signatures of individual training data streams of a plurality of training data streams as locality-sensitive hash (LSH) values associated with the respective training data streams;

determining, based at least in part on the signatures and a predetermined difference criterion associated with the respective LSH values, a training set comprising at least some of the plurality of training data streams and a validation set comprising at least some of the plurality of training data streams;

determining a computational model based at least in part on the training set; and

operating the computational model based at least in part on a trial data stream to provide a trial model output.

2. The method according to claim 1 , wherein the trial model output indicates whether the trial data stream is associated with malware.

3. The method according to claim 1 , wherein at least one of the plurality of training data streams comprises at least part of an executable file.

4. The method according to claim 1 , further comprising:

determining the computational model by performing a supervised learning process using at least one training stream of the training set as training data;

testing the computational model based at least in part on at least one validation stream of the validation set; and

selectively updating the computational model based at least in part on a result of the testing.

5. The method according to claim 1 , wherein:

the method further comprises determining the validation set including validation streams of the plurality of training data streams that satisfy the predetermined difference criterion with respect to training stream(s) in the training set; and

the predetermined difference criterion is defined with respect to the signatures.

6. A method comprising, under control of at least one processing unit:

determining signatures of respective data streams as locality-sensitive hash (LSH) values associated with the respective training data streams;

determining, based at least in part on the signatures, a training set comprising at least one of the data streams and a validation set comprising at least one of the data streams, wherein the respective signatures of a first data stream of the training set and a second data stream of the validation set satisfy a predetermined difference criterion associated with the respective LSH values; and

determining a computational model based at least in part on the training set.

7. The method according to claim 6 , further comprising:

determining respective feature vectors of at least some of the data streams; and

determining the signatures comprising the respective LSH values of the respective feature vectors.

8. The method according to claim 6 , further comprising determining the signatures comprising respective dissimilarity values between the respective data streams and a common reference data stream of the data streams.

9. The method according to claim 6 , further comprising:

determining the computational model further based at least in part on a first hyperparameter value;

operating the computational model based at least in part on at least some of the data streams of the validation set to provide respective model outputs;

determining that the model outputs do not satisfy a predetermined completion criterion; and, in response,

determining a second computational model based at least in part on the training set and a second, different hyperparameter value.

10. The method according to claim 6 , further comprising determining the validation set including individual data streams that satisfy the predetermined difference criterion with respect to at least some of the data streams in the training set.

11. The method according to claim 6 , wherein the training set is disjoint from the validation set.

12. The method according to claim 6 , further comprising:

determining a plurality of partitions of the training set based at least in part on the signatures, wherein each partition of the plurality of partitions comprises at least one of the data streams of the training set;

providing individual partitions of the plurality of partitions to respective computing nodes of a plurality of computing nodes via a communications interface communicatively connected with the processing unit;

receiving respective results from individual computing nodes of the plurality of computing nodes; and

determining the computational model based at least in part on the results.

13. A system comprising:

one or more processors; and

memory communicatively coupled to the one or more processors, the memory storing instructions executable by the one or more processors that, when executed by the one or more processors, cause the system to perform operations including:

determining respective signatures of individual training data streams of a plurality of training data streams as locality-sensitive hash (LSH) values associated with the respective training data streams;

determining, based at least in part on the signatures and a predetermined difference criterion associated with the respective LSH values, a training set comprising at least some of the plurality of training data streams and a validation set comprising at least some of the plurality of training data streams;

determining a computational model based at least in part on the training set; and

operating the computational model based at least in part on a trial data stream to provide a trial model output.

14. The system of claim 13 , wherein the trial model output indicates whether the trial data stream is associated with malware.

15. The system of claim 13 , wherein at least one of the plurality of training data streams comprises at least part of an executable file.

16. The system of claim 13 , wherein the instructions, when executed by the one or more processors, cause the system to perform operations further including:

determining the computational model by performing a supervised learning process using at least one training stream of the training set as training data;

testing the computational model based at least in part on at least one validation stream of the validation set; and

selectively updating the computational model based at least in part on a result of the testing.

17. The system of claim 13 , wherein:

the method further comprises determining the validation set including validation streams of the plurality of training data streams that satisfy the predetermined difference criterion with respect to training stream(s) in the training set; and

the predetermined difference criterion is defined with respect to the signatures.

18. The system of claim 13 , wherein the training set is disjoint from the validation set.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jan 6, 2026
From: FIRST-CITIZENS BANK & TRUST COMPANY
To: CROWDSTRIKE HOLDINGS, INC.; CROWDSTRIKE, INC.
Reel/Frame 074202/0710 →
PATENT SECURITY AGREEMENT Recorded Jan 5, 2021
From: CROWDSTRIKE HOLDINGS, INC.; CROWDSTRIKE, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 054899/0848 →
SECURITY INTEREST Recorded Apr 22, 2019
From: CROWDSTRIKE HOLDINGS, INC.; CROWDSTRIKE, INC.; CROWDSTRIKE SERVICES, INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 048953/0205 →
SECURITY INTEREST Recorded Aug 15, 2017
From: CROWDSTRIKE, INC.
To: SILICON VALLEY BANK
Reel/Frame 043300/0283 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2017
From: KRASSER, SVEN; ELKIND, DAVID; MEYER, BRETT; CRENSHAW, PATRICK
To: CROWDSTRIKE, INC.
Reel/Frame 040934/0272 →
Continuity (1)
Related Publication 20180198800A1 · Jul 12, 2018
Cited By (2)
US 12,346,432 US 12,430,436