IP Library › Granted Patent US 12,367,424
Granted Patent B1
US 12,367,424 · App. 18/818,342 · Granted Jul 22, 2025

Data prefiltering for large scale data classification

Inventors: Olga Gdula (Reading, GB); Felix Schwyzer (Berlin, DE); Calin-Bogdan Miron (Nottingham, GB)
Assignee: CrowdStrike, Inc.
G06N20/00H04L63/1416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,424
App. No.
18/818,342
Granted
Jul 22, 2025
Kind
B1
Abstract

Data prefiltering techniques for large scale data classification are disclosed herein. According to an implementation, a machine learning (ML) model can be trained to classify data elements. The ML model can be applied to a first data volume, resulting in determinations of data elements that belong in a relevant classification. The determined data elements can then be used to configure a prefilter. The prefilter can be applied to a second data volume to identify filtered data elements of types that are similar to the determined data elements. The filtered data elements can be provided to the ML model for classification.

Claims (40)

1. A method, comprising:

classifying, by a trained machine learning classifier, first data elements in a first data volume;

wherein the trained machine learning classifier classifies the first data elements by assigning confidence values representing confidence that the first data elements belong in a classification of data elements;

identifying second data elements based on the confidence values, wherein the second data elements comprise a subset of the first data elements having confidence values that satisfy a threshold confidence value;

subsequent to identifying the second data elements, configuring a prefilter based on the second data elements, so that the prefilter is configured based on previously identified second data elements which were identified at least in part via the trained machine learning classifier;

applying the prefilter to a second data volume, resulting in a reduced second data volume comprising third data elements identified within the second data volume; and

classifying, by the trained machine learning classifier, the reduced second data volume by classifying the third data elements identified within the second data volume, wherein the reduced second data volume is classified by the trained machine learning classifier at a reduced computational expense which is less than a computational expense of classifying the second data volume.

2. The method of claim 1 , further comprising training a machine learning classifier on a training data volume, resulting in the trained machine learning classifier.

3. The method of claim 1 , wherein the first data elements, the second data elements, and the third data elements comprise process trees.

4. The method of claim 1 , wherein the classification of data elements comprises a malicious classification for data elements associated with a malicious process.

5. The method of claim 1 , wherein the prefilter comprises a multistage prefilter configured to use a multistage process to identify the third data elements.

6. The method of claim 1 , wherein the prefilter comprises one or more of at least one allow list or at least one exclusions list.

7. The method of claim 1 , wherein configuring the prefilter comprises performing a statistical analysis of the second data elements in order to identify at least one type of data element included in the second data elements.

8. The method of claim 1 , wherein the second data volume comprises telemetry data received from a security sensor deployed in a remote network.

9. A system, comprising:

a processor, and

at least one memory storing instructions executed by the processor to perform actions including:

applying a prefilter to a second data volume, resulting in a reduced second data volume comprising one or more third data elements identified within the second data volume;

wherein the prefilter is configured according to a process comprising:

classifying, by a trained machine learning classifier, first data elements in a first data volume;

wherein the trained machine learning classifier classifies the first data elements by assigning confidence values representing confidence that the first data elements belong in a classification of data elements;

identifying second data elements based on the confidence values, wherein the second data elements comprise a subset of the first data elements having confidence values that satisfy a threshold confidence value; and

subsequent to identifying the second data elements, configuring the prefilter based on the second data elements so that the prefilter is configured based on previously identified second data elements which were identified at least in part via the trained machine learning classifier, whereby the prefilter is configured to identify the one or more third data elements; and

classifying, by the trained machine learning classifier, the reduced second data volume by classifying the one or more of the third data elements identified within the second data volume, wherein the reduced second data volume is classified by the trained machine learning classifier at a reduced computational expense which is less than a computational expense of classifying the second data volume.

10. The system of claim 9 , wherein the prefilter is configured to identify the one or more third data elements based on a comparison of the one or more third data elements to the second data elements.

11. The system of claim 9 , wherein the first data elements, the second data elements, and the third data elements comprise process trees.

12. The system of claim 9 , wherein the classification of data elements comprises a malicious classification for data elements associated with a malicious process.

13. The system of claim 9 , wherein the prefilter comprises a multistage prefilter configured to use a multistage process to identify the third data elements.

14. The system of claim 9 , wherein the prefilter comprises one or more of at least one allow list or at least one exclusions list.

15. The system of claim 9 , wherein configuring the prefilter comprises performing a statistical analysis of the second data elements in order to identify at least one type of data element included in the second data elements.

16. The system of claim 9 , wherein the second data volume comprises telemetry data received from a security sensor deployed in a remote network.

17. A non-transitory computer-readable storage medium storing computer-readable instructions, that when executed by a processor, cause the processor to perform actions comprising:

receiving telemetry data from a security sensor deployed in a remote network;

prefiltering the telemetry data, resulting in a reduced telemetry data volume;

wherein prefiltering the telemetry data comprises applying a prefilter configured to identify process trees matching one or more process tree types, and

wherein the one or more process tree types are based on identified process trees associated, by a trained machine learning classifier, with at least a threshold confidence of being associated with a malicious process, so that the prefilter is configured based on the identified process trees which were identified at least in part via the trained machine learning classifier; and

classifying process trees included in the reduced telemetry data volume according to confidence of the process trees included in the reduced telemetry data volume being associated with the malicious process, wherein the reduced telemetry data volume is classified at a reduced computational expense which is less than a computational expense of classifying the telemetry data.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the prefilter comprises a multistage prefilter configured to use a multistage process to identify the process trees matching the one or more process tree types.

19. The non-transitory computer-readable storage medium of claim 17 , wherein the prefilter comprises one or more of at least one allow list or at least one exclusions list.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the actions further comprise configuring the prefilter at least in part by classifying, by the trained machine learning classifier, first process trees in a first data volume in order to identify the identified process trees.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2024
From: GDULA, OLGA; SCHWYZER, FELIX; MIRON, CALIN-BOGDAN
To: CROWDSTRIKE, INC.
Reel/Frame 068432/0088 →
References Cited (15)
US 7213260B2 · Judge · 2007 [cited by examiner]
US 20070183655A1 · Konig et al. · 2007 [cited by applicant]
US 20120004893A1 · Vaidyanathan · 2012 [cited by examiner]
US 20160226890A1 · Harang · 2016 [cited by applicant]
US 20200382527A1 · Mitelman · 2020 [cited by examiner]
US 20210406368A1 · Agranonik et al. · 2021 [cited by applicant]
US 20220353284A1 · Voros · 2022 [cited by examiner]
US 20220398491A1 · Khanna · 2022 [cited by examiner]
US 20230344843A1 · Zaytsev · 2023 [cited by examiner]
Thaler, Stefan, Vlado Menkovski, and Milan Petkovic. “Deep learning in information security.” arXiv preprint arXiv:1809.04332 (2018). (Year: 2018). [cited by examiner]
Uzun, Birnur, and Serkan Balli. “Performance evaluation of machine learning algorithms for detecting abnormal data traffic in computer networks.” 2020 5th International Conference on Computer Science and Engineering (UB… [cited by examiner]
Geng, Ye, et al. “An efficient network traffic classification method based on combined feature dimensionality reduction.” 2021 IEEE 21st International Conference on Software Quality, Reliability and Security Companion (… [cited by examiner]
Qiang, Chen Qiang, et al. “Ensemble Method For Net Traffic Classification Based on Deep Learning.” 2021 18th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP). I… [cited by examiner]
Lohani, Anuj, et al. “Static Heuristics Classifiers as Pre-Filter for Malware Target Recognition (MATR).” American Journal of Networks and Communications 4.3 (2015): 44-48. (Year: 2015). [cited by examiner]
Pirscoveanu et al., “Analysis of Malware Behavior: Type Classification using Machine Learning,” in the Proceedings of 2015 International conference on cyber situational awareness, data analytics and assessment (CyberSA)… [cited by applicant]
Cited By (1)
US 12,591,668