IP Library Granted Patent US 12,198,022
Granted Patent B2
US 12,198,022 · App. 17/201,907 · Granted Jan 14, 2025

System and method for training machine learning applications

Inventors: Scott B. Miserendino (Baltimore, MD); Donald D. Steiner (McLean, VA); Ryan V. Peters (Elkridge, MD); Guy B. Fairbanks (Centreville, VA)
Assignee: BluVector, Inc.
G06N20/00G06F18/213G06F21/562G06N3/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,022
App. No.
17/201,907
Granted
Jan 14, 2025
Kind
B2
Abstract

Digital object library management systems and methods for machine learning applications are taught herein. Such a method includes populating a digital object library with a number of machine readable digital objects, modifying the digital objects to include additional machine readable data about the digital objects or other digital objects and the relationships among existing digital objects, generating lists of objects for use in construction and verification of machine learning models used to classify unknown objects into one or more categories, building queries to generate object lists, initiating model generation, in which a machine learning model used to classify unknown objects into one or more categories is generated, initiating model evaluation, storing models, object lists, evaluation results, and associations among these objects, generating a visual display of object metadata, lists, relational information, and evaluation results and running distributable algorithms across the library of digital objects.

Claims (37)

1. A method comprising:

generating, for a plurality of files, metadata indicating one or more properties of the plurality of files;

selecting, based on executing one or more queries that restrict membership in a first portion of the plurality of files based on one or more values of the metadata, the first portion of the plurality of files to train one or more machine learning models to classify the plurality of files as malign or benign;

training, based on the first portion of the plurality of files, the one or more machine learning models; and

determining, based on the one or more machine learning models, that a file from a second portion of the plurality of files is malign.

2. The method of claim 1 , wherein the selecting, based on the one or more values of the metadata, the first portion of the plurality of files, comprises:

determining, for each file of the plurality of files, the one or more values of the metadata; and

selecting the first portion of the plurality of files based on the one or more values of the metadata indicating at least one of: a same file-type, creation on a same date, being from a same source, being from different sources, or being malign or benign.

3. The method of claim 1 , wherein the values of the metadata restrict membership in the first portion of the plurality of files to control training bias in the first portion of the plurality of files.

4. The method of claim 1 , wherein the values of the metadata indicate at least that the first portion of the plurality of files are of a same file-type.

5. The method of claim 1 , wherein the generating the metadata comprises determining at least one feature associated with the plurality of files and a similarity metric indicating a number of occurrences of the at least one feature in the plurality of files, wherein the at least one feature comprises at least one of an n-gram, a header field value, image data, or a file length.

6. The method of claim 1 , wherein the generating the metadata is based on receiving a user input comprising at least a portion of the properties.

7. The method of claim 1 , wherein the plurality of files comprises at least one of a video file, an audio file, a document file, or an executable file.

8. A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

generate, for a plurality of files, metadata indicating one or more properties of the plurality of files;

select, based on executing one or more queries that restrict membership in the first portion of the plurality of files based on one or more values of the metadata, a first portion of the plurality of files to train one or more machine learning models to classify the plurality of files as malign or benign;

train, based on the first portion of the plurality of files, the one or more machine learning models; and

determine, based on the one or more machine learning models, that a file from a second portion of the plurality of files is malign.

9. The device of claim 8 , wherein the selecting, based on the one or more values of the metadata, the first portion of the plurality of files, comprises:

determining, for each file of the plurality of files, the one or more values of the metadata; and

selecting the first portion of the plurality of files based on the one or more values of the metadata indicating at least one of: a same file-type, creation on a same date, being from a same source, being from different sources, or being malign or benign.

10. The device of claim 8 , wherein the values of the metadata restrict membership in the first portion of the plurality of files to control training bias in the first portion of the plurality of files.

11. The device of claim 8 , wherein the values of the metadata indicate at least that the first portion of the plurality of files are of a same file-type.

12. The device of claim 8 , wherein the generating the metadata comprises determining at least one feature associated with the plurality of files and a similarity metric indicating a number of occurrences of the at least one feature in the plurality of files, wherein the at least one feature comprises at least one of an n-gram, a header field value, image data, or a file length.

13. The device of claim 8 , wherein the generating the metadata is based on receiving a user input comprising at least a portion of the properties.

14. The device of claim 8 , wherein the plurality of files comprises at least one of a video file, an audio file, a document file, or an executable file.

15. A non-transitory computer-readable medium storing instructions that, when executed, cause:

generating, for a plurality of files, metadata indicating one or more properties of the plurality of files;

selecting, based on executing one or more queries that restrict membership in the first portion of the plurality of files based on one or more values of the metadata, a first portion of the plurality of files to train one or more machine learning models to classify the plurality of files as malign or benign;

training, based on the first portion of the plurality of files, the one or more machine learning models; and

determining, based on the one or more machine learning models, that a file from a second portion of the plurality of files is malign.

16. The non-transitory computer-readable medium of claim 15 , wherein the selecting, based on the one or more values of the metadata, the first portion of the plurality of files, comprises:

determining, for each file of the plurality of files, the one or more values of the metadata; and

selecting the first portion of the plurality of files based on the one or more values of the metadata indicating at least one of: a same file-type, creation on a same date, being from a same source, being from different sources, or being malign or benign.

17. The non-transitory computer-readable medium of claim 15 , wherein the values of the metadata restrict membership in the first portion of the plurality of files to control training bias in the first portion of the plurality of files.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2021
From: MISERENDINO, SCOTT B.; STEINER, DONALD D.; PETERS, RYAN V.; FAIRBANKS, GUY B.
To: NORTHROP GRUMMAN SYSTEMS CORPORATION
Reel/Frame 055613/0621 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2021
From: NORTHROP GRUMMAN SYSTEMS CORPORATION
To: ACUITY SOLUTIONS CORPORATION
Reel/Frame 055617/0512 →
CHANGE OF NAME Recorded Mar 17, 2021
From: ACUITY SOLUTIONS CORPORATION
To: BLUVECTOR, INC.
Reel/Frame 055617/0727 →
Continuity (2)
Continuation 14635711 · Mar 2, 2015
Related Publication 20210374609A1 · Dec 2, 2021
References Cited (30)
US 8682812B1 · Ranjan · 2014 [cited by applicant]
US 20030051026A1 · Carter et al. · 2003 [cited by applicant]
US 20040088680A1 · Pieper et al. · 2004 [cited by applicant]
US 20060277170A1 · Watry et al. · 2006 [cited by applicant]
US 20090300765A1 · Moskovitch et al. · 2009 [cited by applicant]
US 20100256988A1 · Barnhill et al. · 2010 [cited by applicant]
US 20120191630A1 · Breckenridge et al. · 2012 [cited by applicant]
US 20140046880A1 · Breckenridge et al. · 2014 [cited by applicant]
US 20140090061A1 · Avasarala et al. · 2014 [cited by applicant]
US 20150036919A1 · Bourdev et al. · 2015 [cited by applicant]
JP 2002133389A · 2002 [cited by applicant]
JP 2005182696A · 2005 [cited by applicant]
JP 2006285982A · 2006 [cited by applicant]
JP 2007157058A · 2007 [cited by applicant]
JP 2010092413A · 2010 [cited by applicant]
JP 2011034377A · 2011 [cited by applicant]
JP 2014071493A · 2014 [cited by applicant]
JP 2015079504A · 2015 [cited by applicant]
Rahman, et al., FRAppE: Detecting Malicious Facebook Applications, CoNEXT'12, Dec. 13, 2012, pp. 1-12 (Year: 2012). [cited by examiner]
Alzarooni, Malware Variant Detection, Doctoral Thesis, University College London, Mar. 2012, pp. 1-212 (Year: 2012). [cited by examiner]
Zhang, Fast Algorithms for Burst Detection, Doctoral Thesis, Courant Institute of Mathematical Sciences, Sep. 2006, pp. 1-155 (Year: 2006). [cited by examiner]
Blum, Avrim, et al., “Selection of relevant features and examples in machine learning”, Artifical Intelligence., vol. 97, pp. 245-271 (1997). [cited by applicant]
Chu, et al., Map-Reduce for Machine Learning on Multicore, Advances in neural information processing systems, 19, 2006, pp. 281-288 (Year: 2006). [cited by applicant]
Han, Hui, et al., “Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning”, ICCIC, Part I, LNCS, pp. 878-887 (2005). [cited by applicant]
Huang, Jiayuan, et al., “Correcting Sample Selection Bias by Unlabeled Data”, 8 pages, undated. [cited by applicant]
Laskov, et al., Static Detection of Malicious JavaScript-Bearing PDF Documents, Twenty-Seventh Annual Computer Security Applications Conference, ACSAC 2011, 2011, pp. 373-382 (Year: 2011). [cited by applicant]
Mejia, Carolina, et al., “Supporting Competence upon DotLRN through Personalization”, University of Girona, Insitute of Informatics Applications, Spain, pp. 1-7, undated. [cited by applicant]
Odberg, MultiPerspectives: Object Evolution and Schema Modification Management for Object-Oriented Databases, Doctoral Thesis, Norwegian Institute of Technology, 1995, pp. 1-422 (Year: 1995). [cited by applicant]
Raman, Baranidharan, et al., “Enahncing Learning Using Feature and Example Selection”, Journal of Machine Learning Research, pp. 1-37 (2003). [cited by applicant]
Trevino, et al.,Galgo, An R package for Genetic Algorithm Searches (Customized for Variable Selection in Functional Genomics), School of Biosciences, University of Birmingham, 2006, pp. 1-88. [cited by applicant]