IP Library › Granted Patent US 12,730,891
Granted Patent B1
US 12,730,891 · App. 18/066,885 · Granted Sep 8, 2026

Hybrid scanning using deferral learning

Inventors: Michael James Morais (New York, NY); Marion Marschalek (Portand, OR); Wei Ding (Vancouver, CA); Jeffrey Earl Bickford (Thornton, CO); Baris Coskun (Glen Rock, NJ)
Assignee: Amazon Technologies, Inc.
G06F21/566G06N20/20G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,891
App. No.
18/066,885
Granted
Sep 8, 2026
Kind
B1
Abstract

Systems and methods for performing malware scanning for a service provider network are disclosed. In response to accessing one or more files, file attributes may be determined. Confidence values may be generated based on the file attributes and may be used to select a scan operation. Such scan operations may include a scan operation using a 3 rd party malware scan algorithm or a scan operation using a machine learning generated model. The selection of the scan operation may be performed based on a deferral learning model.

Claims (60)

1 . A system, comprising:

one or more hardware computing devices configured to implement a scanner, wherein the one or more hardware computing devices are configured to:

access one or more files to be scanned for malware;

determine one or more file attributes associated with the one or more files; and

select a scan operation to use to scan the one or more files for malware, based on the determined one or more file attributes, from among a plurality of scan operations comprising:

a) a first scan operation using a third-party scan algorithm for detecting malware; and

b) a second scan operation using a machine learning generated model for detecting malware,

wherein to perform the selection based on the determined one or more file attributes, the one or more hardware computing devices are configured to:

use a deferral learning model to select the scan operation to use, wherein the deferral learning model has been trained to select the scan operation based on learned confidence values associated with the one or more file attributes, wherein the learned confidence values are based on accuracy of malware detection in previous results of the first scan operation for other files having the one or more file attributes or previous results of the second scan operation for the other files having the one or more file attributes.

2 . The system of claim 1 , wherein the one or more hardware computing devices are configured to:

receive training data that has been labeled to indicate malware comprised in the training data;

train a machine learning model for the second scan operation using the training data; and

train the deferral learning model using the training data.

3 . The system of claim 2 , wherein the one or more hardware computing devices are configured to:

generate the learned confidence values based on whether predicted malware detection results of the first or second scan operation match malware detection results indicated in labels in the training data.

4 . The system of claim 1 , wherein the one or more hardware computing devices are configured to:

receive user selection preferences towards false positives or false negatives; and

generate the learned confidence values based, at least in part, on the user selection preferences.

5 . The system of claim 1 , wherein the selection of the scan operation is by default biased towards the first scan operation.

6 . The system of claim 1 , wherein the one or more hardware computing devices are configured to:

provide a file with learned confidence values lower than a threshold to a service, wherein the service analyzes the file based on human action.

7 . One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more processors, implement a scanner and cause the scanner to:

determine one or more file attributes associated with one or more files; and

select a scan operation to use to scan the one or more files, based on the determined one or more file attributes, from among a plurality of scan operations comprising:

a) a first scan operation using a third-party scan algorithm for detecting malware; and

b) a second scan operation using a machine learning generated model for detecting malware,

wherein to perform the selection based on the determined one or more file attributes, the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:

use a deferral learning model to select the scan operation to use, wherein the deferral learning model has been trained to select the scan operation based on learned confidence values associated with the one or more file attributes, wherein the learned confidence values are based on accuracy of malware detection in previous results of the first scan operation for other files having the one or more file attributes or previous results of the second scan operation for the other files having the one or more file attributes.

8 . The one or more non-transitory computer readable storage media of claim 7 , wherein the instructions, when executed on or across the one or more processors, further cause the one or more processors to:

receive training data that has been labeled to indicate malware comprised in the training data;

train a machine learning model for the second scan operation using the training data; and

train the deferral learning model using the training data.

9 . The one or more non-transitory computer readable storage media of claim 8 , wherein the instructions, when executed on or across the one or more processors, further cause the one or more processors to:

generate the learned confidence values based on whether predicted malware detection results of the first or second scan operation match malware detection results indicated in labels in the training data.

10 . The one or more non-transitory computer readable storage media of claim 7 , wherein the instructions, when executed on or across the one or more processors, further cause the one or more processors to:

receive user selection preferences towards false positives or false negatives; and

generate the learned confidence values based, at least in part, on the user selection preferences.

11 . The one or more non-transitory computer readable storage media of claim 7 , wherein the instructions, when executed on or across the one or more processors, further cause the one or more processors to:

generate a response action to a user based, at least in part on a malware detection result of the selected first or second scan operation.

12 . The one or more non-transitory computer readable storage media of claim 11 , wherein the response action comprises a malware detection notification.

13 . The one or more non-transitory computer readable storage media of claim 11 , wherein the response action causes the file with the malware to be removed.

14 . The one or more non-transitory computer readable storage media of claim 11 , wherein the response action causes a source of the malware to be blocked.

15 . A method, comprising:

determining one or more file attributes associated with one or more files; and

selecting a scan operation to use to scan the one or more files, based on the determined one or more attributes, from among a plurality of scan operations comprising:

a) a first scan operation using a third-party scan algorithm for detecting malware; and

b) a second scan operation using a machine learning generated model for detecting malware,

wherein performing the selection based on the determined one or more file attributes comprises:

using a deferral learning model to select the scan operation to use, wherein the deferral learning model has been trained to select the scan operation based on learned confidence values associated with the one or more file attributes, wherein the learned confidence values are based on accuracy of malware detection in previous results of the first scan operation for other files having the one or more file attributes or previous results of the second scan operation for the other files having the one or more file attributes.

16 . The method of claim 15 , further comprising:

receiving training data that has been labeled to indicate malware comprised in the training data;

training a machine learning model for the second scan operation using the training data; and

training the deferral learning model using the training data.

17 . The method of claim 16 , further comprising:

generating the learned confidence values based on whether predicted malware detection results of the first or second scan operation match malware detection results indicated in labels in the training data.

18 . The method of claim 15 , further comprising:

receiving user selection preferences towards false positives or false negatives; and

generating the learned confidence values based, at least in part, on the user selection preferences.

19 . The method of claim 15 , wherein the one or more files comprise one or more Linux-based files.

20 . The method of claim 15 , wherein the first and second scan operations are configured to detect malware including viruses, ransomware, cryptominers, worms, viruses, trojans, bot[net]s, adware, spyware, or rootkits.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: MORAIS, MICHAEL JAMES; MARSCHALEK, MARION; DING, WEI; BICKFORD, JEFFREY EARL; COSKUN, BARIS
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062160/0652 →
References Cited (81)
US 7849507B1 · Bloch · 2010 [cited by examiner]
US 8832832B1 · Visbal · 2014 [cited by applicant]
US 9231965B1 · Vasseur · 2016 [cited by applicant]
US 9363282B1 · Yu · 2016 [cited by applicant]
US 9756070B1 · Crowell · 2017 [cited by examiner]
US 10320813B1 · Ahmed et al. · 2019 [cited by applicant]
US 10826933B1 · Ismael · 2020 [cited by applicant]
US 11727113B1 · Briliauskas · 2023 [cited by examiner]
US 12118095B1 · Millar · 2024 [cited by examiner]
US 20030051026A1 · Carter et al. · 2003 [cited by applicant]
US 20040015719A1 · Lee et al. · 2004 [cited by applicant]
US 20050144480A1 · Kim et al. · 2005 [cited by applicant]
US 20070074272A1 · Watanabe · 2007 [cited by applicant]
US 20070094491A1 · Teo et al. · 2007 [cited by applicant]
US 20080033672A1 · Gulati · 2008 [cited by applicant]
US 20080098476A1 · Syversen · 2008 [cited by applicant]
US 20090265778A1 · Wahl · 2009 [cited by applicant]
US 20100007489A1 · Misra et al. · 2010 [cited by applicant]
US 20110023114A1 · Diab · 2011 [cited by applicant]
US 20110060956A1 · Goldsmith · 2011 [cited by applicant]
US 20110214157A1 · Korsunsky · 2011 [cited by applicant]
US 20110225644A1 · Pullikottil et al. · 2011 [cited by applicant]
US 20120084859A1 · Radinsky · 2012 [cited by examiner]
US 20120110667A1 · Zubrilin · 2012 [cited by examiner]
US 20120284793A1 · Steinbrecher · 2012 [cited by applicant]
US 20140059683A1 · Ashley · 2014 [cited by applicant]
US 20140223555A1 · Sanz Hernando et al. · 2014 [cited by applicant]
US 20140289856A1 · Jiang et al. · 2014 [cited by applicant]
US 20140380466A1 · Schultz · 2014 [cited by applicant]
US 20150033341A1 · Schmidtier · 2015 [cited by applicant]
US 20150067857A1 · Symons · 2015 [cited by applicant]
US 20150304343A1 · Cabrera · 2015 [cited by applicant]
US 20150355957A1 · Steiner · 2015 [cited by applicant]
US 20160028753A1 · Di Pietro · 2016 [cited by applicant]
US 20160028754A1 · Crus · 2016 [cited by applicant]
US 20160078362A1 · Christodorescu · 2016 [cited by applicant]
US 20160087861A1 · Kuan · 2016 [cited by applicant]
US 20160099963A1 · Mahaffey · 2016 [cited by applicant]
US 20160164886A1 · Thrash · 2016 [cited by applicant]
US 20160191545A1 · Nanda · 2016 [cited by applicant]
US 20160212012A1 · Young · 2016 [cited by applicant]
US 20170063891A1 · Muddu · 2017 [cited by applicant]
US 20170070528A1 · Coskun · 2017 [cited by applicant]
US 20170134397A1 · Dennison · 2017 [cited by examiner]
US 20170262633A1 · Miserendino · 2017 [cited by examiner]
US 20170300693A1 · Zhang · 2017 [cited by examiner]
US 20180026995A1 · Dufour · 2018 [cited by examiner]
US 20180027006A1 · Zimmermann · 2018 [cited by applicant]
US 20180082064A1 · Wang · 2018 [cited by examiner]
US 20180103056A1 · Kohout · 2018 [cited by applicant]
US 20190132787A1 · Ryan · 2019 [cited by examiner]
US 20190138938A1 · Vasseur · 2019 [cited by applicant]
US 20190297096A1 · Ahmed et al. · 2019 [cited by applicant]
US 20200380160A1 · Kraus · 2020 [cited by examiner]
US 20210281592A1 · Givental · 2021 [cited by examiner]
US 20220036208A1 · Rao · 2022 [cited by examiner]
US 20220230070A1 · Shabtai · 2022 [cited by examiner]
US 20220414214A1 · Strogov · 2022 [cited by examiner]
US 20230214485A1 · Beek · 2023 [cited by examiner]
US 20230229782A1 · Mohanty · 2023 [cited by examiner]
Fire Eye Inc., “The Business Case for Protecting Against Advanced Attacks”, https://www2.fireeye.com/StopTheNoise-IDC-Numbers-Game-Special-Report.html., dated 2014, pp. 1-13. [cited by applicant]
S. Mathew, D. Britt, R. Giomundo, S. Upadhyaya, M. Sudit, and A. Stotz. Realtime multistage attack awareness through enhanced intrusion alert clustering. In MILCOM 2005—2005 IEEE Military Communications Conference, vol.… [cited by applicant]
S. Mathew, C. Shah, and S. Upadhyaya. An alert fusion framework for situation awareness of coordinated multistage attacks. In Third IEEE International Workshop on Information Assurance (IWIA'05), pp. 95-104, Mar. 2005. [cited by applicant]
Wajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen, Kangkook Jee, Zhichun Li, and Adam Bates. Nodoze: Combatting threat alert fatigue with automated provenance triage. In NDSS, 2019. [cited by applicant]
Steven Noel, Eric Harley, Kam Him Tam, and Greg Gyor. Big-data architecture for cyber attack graphs representing security relationships in nosql graph databases. 2014. [cited by applicant]
S. Noel, P. D. Rowe, S. Purdy, M. Limiero, T. Lu, and W. Mathews. Missionfocused cyber situational understanding via graph analytics. In 2018 10th International Conference on Cyber Conflict (CyCon), pp. 427-448, 2018. [cited by applicant]
Wikipedia contributors. Sqrrl—Wikipedia, the free encyclopedia. https://en.wikipedia.org/w/index.php?title=Sqrrl&oldid=899580655, 2022. [cited by applicant]
Wikipedia contributors. Birthday problem—Wikipedia, the free encyclopedia. https://en.wikipedia.org/w/index.php?title=Birthday_problem&oldid=912594355, 2022. [cited by applicant]
Matei Zaharia, Reynold S. Xin, PatrickWendell, Tathagata Das, Michael Armbrust, Ankur Dave, Xiangrui Meng, Josh Rosen, Shivaram Venkataraman, Michael J. Franklin, Ali Ghodsi, Joseph Gonzalez, Scott Shenker, and Ion Stoi… [cited by applicant]
Joseph E. Gonzalez, Reynold S. Xin, Ankur Dave, Daniel Crankshaw, Michael J. Franklin, and Ion Stoica. Graphx: Graph processing in a distributed dataflow framework. In Proceedings of the 11th USENIX Conference on Operat… [cited by applicant]
Quissem Ben Fredj. A realistic graph-based alert correlation system. Security and Communication Networks, 8 (15):2477-2493, 2015. [cited by applicant]
U.S. Appl. No. 17/809,519, filed Jun. 28, 2022, McCubbin, et al. [cited by applicant]
U.S. Appl. No. 18/065,481, filed Dec. 13, 2022, Zhang et al. [cited by applicant]
AWS, “Amazon GuardDuty,” downloaded from https://aws.amazon.com/guardduty/ on Dec. 20, 2022, pp. 1-8. [cited by applicant]
AWS, “AWS CloudTrail,” downloaded from https://aws.amazon.com/cloudtrail/ on Dec. 20, 2022, pp. 1-7. [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention is all you need.” In Advances in neural information processing systems, pp. 5998-60… [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, Version 2 2019, pp. 1-16. [cited by applicant]
Min Du, Feifei Li, Guineng Zheng, and Vivek Srikumar. “Deeplog: Anomaly detection and diagnosis from system logs through deep learning.” In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Se… [cited by applicant]
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. “Deep learning”. MIT press, Oct. 3, 2015, pp. 1-705. [cited by applicant]
Min-hwan Oh and Garud Iyengar. “Sequential anomaly detection using inverse reinforcement learning.” In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1480-1490, 201… [cited by applicant]
Noveen Sachdeva, Giuseppe Manco, Ettore Ritacco, and Vikram Pudi. “Sequential variational autoencoders for collaborative filtering”. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mini… [cited by applicant]