IP Library Granted Patent US 12,524,530
Granted Patent B2
US 12,524,530 · App. 18/089,373 · Granted Jan 13, 2026

Attack detection and countermeasure identification system

Inventors: Jean-Paul Watson (Livermore, CA); Alyson Lindsey Fox (San Leandro, CA); Sarah Camille Mousley Mackay (Livermore, CA); Wayne Bradford Mitchell (Hayward, CA); Matthew Landen (Atlanta, GA); Key-Whan Chung (Urbana, IL); Elizabeth Diane Reed (Urbana, IL)
Assignee: Lawrence Livermore National Security
G06F21/554G06N3/092G06F2221/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,530
App. No.
18/089,373
Granted
Jan 13, 2026
Kind
B2
Abstract

A method is disclosed which comprises accessing a detector model that is trained in parallel with an operator model and an attacker model using a reinforcement learning technique based on iteratively simulating scenarios of operation of an environment to generate training data and learning weights of the models based on the simulated training data. The simulating of a scenario is based on the last learned weights of the models. The method further comprises, during operation of the environment, applying the detector model to an operator action, a prior observation of state of the environment from prior to taking the operator action, and a current observation of the environment from after taking the operator action, to detect whether an attack on the environment has occurred.

Claims (39)

1 . A method performed by one or more computing systems to support responding to a cyberattack on a physical infrastructure system via a computer network environment, the method comprising:

accessing a specification of the physical infrastructure system that includes components having a plurality of states;

running scenarios, comprising virtual simulations of the physical infrastructure system that output machine learning model training data, to modify a current state of the physical infrastructure system corresponding to the plurality of states, wherein running a scenario includes:

modifying the current state of the physical infrastructure system based on an operator action, wherein a modification to the plurality of states includes a simulated change to a physical infrastructure topology;

modifying the modified current state of the physical infrastructure system based on an attacker action to generate a new state; and

detecting within the scenario whether an attack on the physical infrastructure system has occurred based on the operator action, the current state, and the new state;

and

training an operator model and a detector model based on the operator action, the attacker action, and a detection of whether an attack on the physical infrastructure system has occurred, wherein the operator model is trained to identify an effective operator action given a particular current state of the physical infrastructure system and the detector model is trained to detect an attack on the physical infrastructure system and said training modifies weights assigned to the operation action and the attacker action as associated with the particular current state and the detection of the attack respectively.

2 . The method of claim 1 , further comprising:

training an attacker model in parallel with training the operator model and the detector model, based on the operator actions, the attacker actions, and the detections of the scenarios, wherein the attacker model is trained to identify effective attacks on the physical infrastructure system.

3 . The method of claim 1 , wherein the running a scenario generates an operator reward for each operator action as an indication of effectiveness of the operator action, an attacker reward for each attacker action as an indication of effectiveness of the attacker action, and a detector reward as an indication of effectiveness of the detection, and wherein the training the operator model and the detector model step factors in the operator reward, the attacker reward, and the detector reward.

4 . The method of claim 1 , further comprising:

receiving a current state of a non-simulated, real environment of a physical infrastructure system, an operator action to modify the current state, and a new state after modification of the current state; and

applying the detector model to the operator action, the current state, and the new state to detect whether an attack has occurred on the non-simulated, real environment of the physical infrastructure system.

5 . The method of claim 4 , further comprising:

applying the operator model to identify an effective operator action when an attack is detected.

6 . The method of claim 1 , wherein the running scenarios and the training are performed iteratively, wherein the running employs the operator model, an attacker model, and the detector model that was last trained, respectively, to generate operator actions, to generate an attacker action, and to detect an attack.

7 . The method of claim 1 , wherein the physical infrastructure system is a power grid system includes generators, loads, substations, and lines.

8 . The method of claim 1 , wherein the computer network environment is an information technology (IT) environment.

9 . A processing system configured to support responding to a cyberattack on a physical infrastructure system via a computer network environment, the system comprising:

at least one processor; and

at least one non-transitory computer-readable storage medium storing instructions, execution of which by the at least one processor causes the processing system to perform operations comprising:

accessing a specification of the physical infrastructure system that includes components having a plurality of states;

running scenarios, comprising virtual simulations of the physical infrastructure system that output machine learning model training data, to modify a current state of the physical infrastructure system corresponding to the plurality of states, wherein running a scenario includes:

modifying the current state of the physical infrastructure system based on an operator action, wherein a modification to the plurality of states includes a simulated change to a physical infrastructure topology;

modifying the modified current state of the physical infrastructure system based on an attacker action to generate a new state; and

detecting within the scenario whether an attack on the physical infrastructure system has occurred based on the operator action, the current state, and the new state; and

training an operator model and a detector model based on the operator action, the attacker action, and a detection of whether an attack on the physical infrastructure system has occurred, wherein the operator model is trained to identify an effective operator action given a particular current state of the physical infrastructure system and the detector model is trained to detect an attack on the physical infrastructure system and said training modifies weights assigned to the operation action and the attacker action as associated with the particular current state and the detection of the attack respectively.

10 . The system of claim 9 , the operations further comprising:

training an attacker model in parallel with training the operator model and the detector model, based on the operator actions, the attacker actions, and the detections of the scenarios, wherein the attacker model is trained to identify effective attacks on the physical infrastructure system.

11 . The system of claim 9 , wherein the running a scenario generates an operator reward for each operator action as an indication of effectiveness of the operator action, an attacker reward for each attacker action as an indication of effectiveness of the attacker action, and a detector reward as an indication of effectiveness of the detection, and wherein the training the operator model and the detector model step factors in the operator reward, the attacker reward, and the detector reward.

12 . The system of claim 9 , the operations further comprising:

receiving a current state of a non-simulated, real environment of a physical infrastructure system, an operator action to modify the current state, and a new state after modification of the current state; and

applying the detector model to the operator action, the current state, and the new state to detect whether an attack has occurred on the non-simulated, real environment of the physical infrastructure system.

13 . The system of claim 12 , the operations further comprising:

applying the operator model to identify an effective operator action when an attack is detected.

14 . The system of claim 9 , wherein the running scenarios and the training are performed iteratively, wherein the running employs the operator model, an attacker model, and the detector model that was last trained, respectively, to generate operator actions, to generate an attacker action, and to detect an attack.

15 . The system of claim 9 , wherein the physical infrastructure system is a power grid system and the components include generators, loads, substations, and lines.

16 . The system of claim 9 , wherein the computer network environment is an information technology (IT) environment.

Assignments (2)
CONFIRMATORY LICENSE (SEE DOCUMENT FOR DETAILS) Recorded Jan 26, 2023
From: LAWRENCE LIVERMORE NATIONAL SECURITY, LLC
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 062513/0603 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2023
From: WATSON, JEAN-PAUL; FOX, ALYSON LINDSEY; MACKAY, SARAH CAMILLE MOUSLEY; MITCHELL, WAYNE BRADFORD; LANDEN, MATTHEW; CHUNG, KEY-WHAN; REED, ELIZABETH DIANE
To: LAWRENCE LIVERMORE NATIONAL SECURITY, LLC
Reel/Frame 062276/0536 →
Continuity (2)
Provisional Application 63294064 · Dec 27, 2021
Related Publication 20230297672A1 · Sep 21, 2023
References Cited (138)
US 9043894B1 · Dennison · 2015 [cited by examiner]
US 9779271B2 · Kanakarajan · 2017 [cited by examiner]
US 10452845B2 · Mestha · 2019 [cited by examiner]
US 10581235B2 · Seuss · 2020 [cited by examiner]
US 10601862B1 · Kurupati · 2020 [cited by examiner]
US 11050234B2 · Schweitzer, III · 2021 [cited by examiner]
US 11165812B2 · Satish · 2021 [cited by examiner]
US 11665194B2 · Reddy · 2023 [cited by examiner]
US 11689021B2 · Wu · 2023 [cited by examiner]
US 11956255B1 · Holub · 2024 [cited by examiner]
US 11972335B2 · Dasgupta · 2024 [cited by examiner]
US 12113810B2 · Hariri · 2024 [cited by examiner]
US 12184697B2 · Crabtree · 2024 [cited by examiner]
US 12200014B2 · El Gamal · 2025 [cited by examiner]
US 20070198235A1 · Takeuchi · 2007 [cited by examiner]
US 20130132149A1 · Wei · 2013 [cited by examiner]
US 20150163242A1 · Laidlaw · 2015 [cited by examiner]
US 20160028753A1 · Di Pietro · 2016 [cited by examiner]
US 20170093889A1 · McEachern · 2017 [cited by examiner]
US 20170104780A1 · Zaffarano · 2017 [cited by examiner]
US 20170279839A1 · Vasseur · 2017 [cited by examiner]
US 20180024900A1 · Premerlani · 2018 [cited by examiner]
US 20180260561A1 · Mestha · 2018 [cited by examiner]
US 20180262525A1 · Yan · 2018 [cited by examiner]
US 20190260196A1 · Seuss · 2019 [cited by examiner]
US 20190260768A1 · Mestha · 2019 [cited by examiner]
US 20200119556A1 · Shi · 2020 [cited by examiner]
US 20200228566A1 · Kurupati · 2020 [cited by examiner]
US 20200287930A1 · Satish · 2020 [cited by examiner]
US 20200327411A1 · Shi · 2020 [cited by examiner]
US 20210051162A1 · Taylor · 2021 [cited by examiner]
US 20210158162A1 · Hafner · 2021 [cited by examiner]
US 20210182385A1 · Roychowdhury · 2021 [cited by examiner]
US 20210201156A1 · Hafner · 2021 [cited by examiner]
US 20210264795A1 · Mguni · 2021 [cited by examiner]
US 20210295176A1 · Jacobs · 2021 [cited by examiner]
US 20210334441A1 · Wu · 2021 [cited by examiner]
US 20210356923A1 · Wu · 2021 [cited by examiner]
US 20220019674A1 · Frey · 2022 [cited by examiner]
US 20220036186A1 · Refaat · 2022 [cited by examiner]
US 20220131366A1 · Schweitzer, III · 2022 [cited by examiner]
US 20220173591A1 · Cova Acosta · 2022 [cited by examiner]
US 20220201042A1 · Crabtree · 2022 [cited by examiner]
US 20220210200A1 · Crabtree · 2022 [cited by examiner]
US 20220224723A1 · Crabtree · 2022 [cited by examiner]
US 20220245441A1 · Dechene · 2022 [cited by examiner]
US 20220284156A1 · Zheng · 2022 [cited by examiner]
US 20220303290A1 · Baidya · 2022 [cited by examiner]
US 20220343117A1 · Jeong · 2022 [cited by examiner]
US 20220343230A1 · Casey · 2022 [cited by examiner]
US 20220345479A1 · Markonis · 2022 [cited by examiner]
US 20220357729A1 · Xu · 2022 [cited by examiner]
US 20220360597A1 · Fellows · 2022 [cited by examiner]
US 20230073326A1 · Schrittwieser · 2023 [cited by examiner]
US 20230115046A1 · Karta · 2023 [cited by examiner]
US 20230177165A1 · Underwood · 2023 [cited by examiner]
US 20230185912A1 · Sinn · 2023 [cited by examiner]
US 20230325511A1 · Jaster · 2023 [cited by examiner]
US 20230378753A1 · Marinakis · 2023 [cited by examiner]
US 20240061939A1 · Pieczul · 2024 [cited by examiner]
US 20240119298A1 · Pham · 2024 [cited by examiner]
US 20240176894A1 · Lev · 2024 [cited by examiner]
US 20240311639A1 · Hansen · 2024 [cited by examiner]
US 20250023902A1 · Thepie Fapi · 2025 [cited by examiner]
P. Xu et al., “Active Power Correction Strategies Based on Deep Reinforcement Learning—Part I: A Simulation-driven Solution for Robustness,”in CSEE Journal of Power and Energy Systems, vol. 8, No. 4, pp. 1122-1133, Jul.… [cited by examiner]
Chen et al., Understanding the Safety Requirements for Learning-based Power Systems Operations, arXiv,Oct. 2021. [cited by examiner]
D. An, Q. Yang, W. Liu, and Y. Zhang, “Defending Against Data Integrity Attacks in Smart Grid: A Deep Reinforcement Learning-Based Approach,” in IEEE Access, vol. 7, pp. 110835-110845, 2019. [cited by examiner]
S. Paul, Z.Ni and C.Mu, “A Learning-Based Solution for an Adversarial Repeated Game in Cyber—Physical Power Systems,” in IEEE Transactions on Neural Networks and Learning Systems, vol. 31, No. 11, pp. 4512-4523, Nov. 20… [cited by examiner]
L. Omnes, A. Marot, and B.Donnot, “Adversarial Training for a Continuous Robustness Control Problem in Power Systems,” 2021 IEEE Madrid Power Tech, Madrid, Spain, 2021. [cited by examiner]
An, Dou, et al., “Defending against data integrity attacks in smart grid: A deep reinforcement learning-based approach,”. IEEE Access, 7:110835-110845, 2019. [cited by applicant]
Blackenergy apt attacks in ukraine. https://www.kaspersky.com/resource-center/threats/blackenergy. [cited by applicant]
Chen Binbin. AsprinChina/L2RPN_nips_2020_a_ppo_solution, May 2021. original-date: 2020-11-12T03:53:04Z. Accessed via: https://github.com/AsprinChina/L2RPN_NIPS_2020_a_PPO_Solution. [cited by applicant]
Chung, Keywhan, et al., “Game theory with learning for cyber security monitoring,” In 2016 IEEE 17th International Symposium on High Assurance Systems Engineering (HASE), pp. 1-8. IEEE, 2016. [cited by applicant]
Clemente, Alfredo V., et al., “Efficient parallel methods for deep reinforcement learning,” arXiv preprint arXiv:1705.04862, 2017. [cited by applicant]
Dán, György, et al., “Stealth attacks and protection schemes for state estimators in power systems,” In 2010 first IEEE international conference on smart grid communications, pp. 214-219. IEEE, 2010. [cited by applicant]
Daumé, Hal, “A course in machine learning,” 2015. [cited by applicant]
Deokar, Bhagyashree, et al., “Intrusion detection system using log files and reinforcement learning,” International Journal of Computer Applications, 45(19):28-35, 2012. [cited by applicant]
Dragonfly: Western energy sector targeted by sophisticated attack group, accessed via: https://symantec-enterprise-blogs.security.com/blogs/threat-intelligence/dragonfly-energy-sector-cyber-attacks. [cited by applicant]
Dragos, Inc., “2020 ICS cybersecurity year in review,” Apr. 2021. Accessed Via: https://www.dragos.com/blog/industry-news/2020-ics-cybersecurity-year-in-review/. [cited by applicant]
Dragos, Inc., “The evolution of cyber attacks on electric operations,” 2019. Accessed via: https://www.dragos.com/blog/industry-news/the-evolution-of-cyber-attacks-on-electric-operations/. [cited by applicant]
Eremia, Mircea, et al., “Handbook of electrical power system dynamics: modeling, stability, and control,” vol. 92. John Wiley & Sons, 2013. [cited by applicant]
Falcon complete: Managed detection and response, accessed via: https://www.crowdstrike.com/endpoint-security-products/falcon-complete/. [cited by applicant]
Ferdowsi, Aidin, et al., “Robust deep reinforcement learning for security and safety in autonomous vehicle systems,” In 2018 21st International Conference on Intelligent Transportation Systems (ITSC), pp. 307-312. IEEE,… [cited by applicant]
Garcia, Luis A., et al., “Hey, my malware knows physics! Attacking plcs with physical model aware rootkit,” In NDSS, 2017. [cited by applicant]
Github, Grid2op: a testbed platform to model sequential decision making in power systems. Accessed via: https://github.com/rte-france/Grid2Op. [cited by applicant]
Github, Lujasone. Lujasone/neurips_2020_l2rpn_comp_an _approach: The implementation of neurips_2020_12rpn_track1 (robustness) and track2 (adaptability) competition. Accessed via: https://github.com/lujasone/NeurIPS_2020… [cited by applicant]
Glenn, Colleen, et al., “Cyber threat and vulnerability analysis of the U.S. electric sector,” Technical report, Idaho National Lab. (INL), Idaho Falls, ID (United States), 2016. [cited by applicant]
Glover, J Duncan, et al., “Power system analysis & design,” SI version. Cen gage Learning, 2012. [cited by applicant]
Gupta, Abhishek, et al., “Adversarial reinforcement learning for observer design in autonomous systems under cyber attacks,” arXiv preprint arXiv:1809.06784, 2018. [cited by applicant]
György, András, et al., “Efficient multi-start strategies for local search algorithms,” Journal of Artificial Intelligence Research, 41:407-444, 2011. [cited by applicant]
Han, Yi, et al., “Reinforcement learning for autonomous defence in software-defined networking,” In International Conference on Decision and Game Theory for Security, pp. 145-165. Springer, 2018. [cited by applicant]
Hausknecht, Matthew, et al., “Deep recurrent q-learning for partially observable mdps,” In 2015 aaai fall symposium series, 2015. [cited by applicant]
Illinois Information Trust Institute, Illinois Center for a Smarter Electric Grid (ICSEG), IEEE 14-bus system. Accessed via: https://icseg.iti.illinois.edu/ieee-14-bus-system. [cited by applicant]
Kaelbling, Leslie Pack, et al., “Planning and acting in partially observable stochastic domains,” Artificial Intelligence, 101(1):99-134, 1998. [cited by applicant]
Kurt, Mehmet Necip, et al., “Online cyber-attack detection in smart grid: A reinforcement learning approach,” IEEE Trans actions on Smart Grid, 10(5):5174-5185, 2019. [cited by applicant]
Learning to Run a Power Network—Neurips Track 1, L2rpn neurips 2020—robustness track. Accessed via: https://competitions.codalab.org/competitions/25426. [cited by applicant]
Lee, Robert M, et al., “Crashoverride: Analysis of the threat to electric grid operations,” Dragos Inc., Mar. 2017. [cited by applicant]
Li, Yanda, et al., “Mobile cloud offloading for malware detections with learning,” In 2015 IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS), pp. 197-201. IEEE, 2015. [cited by applicant]
Li, Yuancheng, et al., “False data injection attacks with incomplete network topology information in smart grid,” IEEE Access, 7:3656-3664, 2018. [cited by applicant]
Liu, Yao, et al., “False data injection attacks against state estimation in electric power grids,” ACM Transactions on Information and System Security (TISSEC), 14(1):1-33, 2011. [cited by applicant]
López-Morales, Efrén, et al., “Honeyplc: A next-generation hon eypot for industrial control systems,” In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, pp. 279-291, 2020. [cited by applicant]
Malialis, Kleanthis, et al., “Distributed response to network intrusions using multiagent reinforcement learning,” Engineering Applications of Artificial Intelligence, 41:270-284, 2015. [cited by applicant]
Mandlekar, Ajay, et al., “Adversarially robust policy learning: Active construction of physically-plausible perturbations,” In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 3932-39… [cited by applicant]
Marot, Antoine, et al., “Learning to run a power network challenge for training topology controllers,” Electric Power Systems Research, 189:106635, 2020. [cited by applicant]
Marot, Antoine, et al., “Learning to run a power network challenge: a retrospective analysis,” arXiv preprint arXiv:2103.03104, 2021. [cited by applicant]
Minih, Volodymyr, et al., “Asynchronous methods for deep reinforcement learning”, In International conference on machine learning, pp. 1928-1937. PMLR, 2016. [cited by applicant]
Nair, Arun, et al. “Massively parallel methods for deep reinforcement learning,” arXiv preprint arXiv:1507.04296, 2015. [cited by applicant]
Next-generation firewalls. Accessed via: https://www.paloaltonetworks.com/network-security/next-generation-firewall. [cited by applicant]
Nguyen, Thanh Thi, et al., “Deep reinforcement learning for cyber security,” arXiv preprint arXiv:1906.05799, 2021. [cited by applicant]
Omnes, Loïc, et al., “Adversarial training for continuous robustness control problem in power systems,” arXiv preprint arXiv:2012.11390, 2020. [cited by applicant]
Paszke, Adam, et el., “Pytorch: An imperative style, high-performance deep learning library,” Advances in Neural Information Processing Systems 32, pp. 8024-8035. Curran Associates, Inc., 2019. [cited by applicant]
Pateria, Shubham, et al., “Hierarchical reinforcement learning: A comprehensive survey,” ACM Computing Surveys (CSUR), 54(5):1-35, 2021. [cited by applicant]
Powers, David MW, “Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation,” arXiv preprint arXiv:2010.16061, 2020. [cited by applicant]
Shao, Wei, et al., “Corrective switching algorithm for relieving overloads and voltage violations,” IEEE Transactions on Power Systems, 20(4):1877-1885, 2005. [cited by applicant]
Shekari, Tohid, et al., “Rfdids: Radio frequency-based distributed intrusion detection system for the power grid,” In NDSS, 2019. [cited by applicant]
Slowik, Joe, “Crashoverride: Reassessing the 2016 ukraine electric power event as a protection-focused attack,” Dragos, Inc, 2019. [cited by applicant]
Slowik, Joe, “Stuxnet to crashoverride to trisis: Evaluating the history & future of integrity-based attacks on industrial environments,” 2020. [cited by applicant]
Soltan, Saleh, et al., “Blackiot: lot botnet of high wattage devices can disrupt the power grid,” In 27th fUSENIXg Security Symposium (fUSENIXg Security 18), pp. 15-32, 2018. [cited by applicant]
Stooke, Adam, et al., “Accelerated methods for deep reinforcement learning,” arXiv preprint arXiv:1803.02811, 2019. [cited by applicant]
Sutton, Richard S, et al., “Reinforcement learning: An introduction,” MIT press, 2018. [cited by applicant]
Tian, Jiwei, et al., “Data-driven and low-sparsity false data injection attacks in smart grid,” Security and Communication Networks, 2018. [cited by applicant]
Trellix Email Security, accessed via: https://www.trellix.com/en-us/products/email-security.html. [cited by applicant]
Trellix Endpoint security software and solutions. Accessed via: https://www.fireeye.com/products/endpoint-security.html. [cited by applicant]
Van Hasselt, Hado, et al., “Deep reinforcement learning with double q-learning.,” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, 2016. [cited by applicant]
Wan, Xiaoyue, et al., “Reinforcement learning based mobile offloading for cloud-based malware detection,” In GLOBECOM 2017-2017 IEEE Global Communications Conference, pp. 1-6. IEEE, 2017. [cited by applicant]
Wang, Zhenhua, et al., “Coordinated topology attacks in smart grid using deep reinforcement learning,” IEEE Transactions on Industrial Informatics, 17(2):1407-1415, 2021. [cited by applicant]
Wikipedia, “Reinforcement learning,” Retrieved from: Retrieved from “https://en.wikipedia.org/w/index.php?title=Reinforcement_learning&oldid=1046562875”. [cited by applicant]
Wilson, David, et al., “Deep learning-aided cyber-attack detection in power transmission systems,” In 2018 IEEE Power & Energy Society General Meeting (PESGM), pp. 1-5. IEEE, 2018. [cited by applicant]
Wood, Allen J., et al., “Power generation, operation, and control,” John Wiley & Sons, 2013. [cited by applicant]
Wu, Xian, et al., “Adversarial policy training against deep reinforcement learning,” In 30th fUSENIXg Security Symposium (fUSENIXg Security 21), 2021. [cited by applicant]
Xiao, Liang, et al., “Spoofing detection with reinforcement learning in wireless networks,” In 2015 IEEE Global Communications Conference (GLOBECOM), pp. 1-5. IEEE, 2015. [cited by applicant]
Xu, Xin, “Sequential anomaly detection based on temporal-difference learning: Principles, models and case studies,” Applied Soft Computing, 10(3):859-867, 2010. [cited by applicant]
Xu, Xin, et al., “A kernel-based reinforcement learning approach to dynamic behavior modeling of intrusion detection,” In International Symposium on Neural Networks, pp. 455-464. Springer, 2007. [cited by applicant]
Xu, Xin, et al., “A reinforcement learning approach for host-based intrusion detection using sequences of system calls,” In International Conference on Intelligent Computing, pp. 995-1003. Springer, 2005. [cited by applicant]
Yau, David KY, et al., “Defending against distributed denial-of-service attacks with max-min fair server-centric router throttles,” IEEE/ACM Transactions on Networking, 13(1):29-42, 2005. [cited by applicant]
Zhang, Chiyuan, et al., “A study on overfitting in deep reinforcement learning,” arXiv preprint arXiv:1804.06893, 2018. [cited by applicant]
Zhou, Bo, et al., “Action set based policy optimization for safe power grid management,” 2021. [cited by applicant]
Zhu, Minghui, et al., “Reinforcement learning algorithms for adaptive cyber defense against heartbleed,” In Proceedings of the First ACM Workshop on Moving Target Defense, pp. 51-58, 2014. [cited by applicant]