IP Library Granted Patent US 12675727
Granted Patent B2
US 12675727 · App. 16/998,680 · Granted Jul 7, 2026

Method and system for determining policies, rules, and agent characteristics, for automating agents, and protection

Inventors: Ulrich Lang (San Diego, CA); Rudolf Schreiner (Falkensee, DE)
Assignee: ObjectSecurity LLC
G06N20/00G06N3/092G06N5/025G06N5/04G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675727
App. No.
16/998,680
Granted
Jul 7, 2026
Kind
B2
Abstract

A method of automatically configuring an action determination model includes determining an environment model, determining an action determination model that indicates an action option, determining whether the action determination model indicates a next action option, and if so, determining an action based on the action determination model, simulating execution of the action across the environment model, obtaining a simulated result, adjusting the action determination model. Then, until environment or an agent reach an end state, the following are repeated: determining whether the action determination model indicates the next action option, and if so, determining the action based on the action determination model, simulating the execution of the action across the environment model, obtaining the simulated result, and adjusting the action determination model.

Claims (14)

1 . A computer-implemented method of automatically configuring at least one action determination model for an agent device of at least one agent connected to at least one simulated Information Technology (IT) environment via a computer network, the agent device of the at least one agent having the at least one action determination model indicating at least one action option to be simulated for execution by the agent device of the at least one agent to the at least one simulated IT environment, the method comprising:

generating and configuring, via a processor, the at least one simulated IT environment comprising a plurality of computer-virtualized and/or real networked IT systems interconnected via a defined network topology for performing IT-system operations, wherein the at least one simulated IT environment models hardware and software components of real-world IT systems;

storing the at least one action determination model in a storage device, each action determination model including a plurality of machine-learning (ML) based artificial intelligent (AI) neural net models each comprising an IT-system operation and defining at least one action to be simulated for execution by the agent device of the at least one agent to the at least one simulated IT environment in at least one given context of attributes, state and/or behaviors of the at least one simulated IT environment for performing the IT-system operation by the IT systems in the at least one simulated IT environment and having a different set of weighted parameters for a corresponding one of the at least one action, the IT-system operation modifying a state of at least one component within the simulated IT environment, including system configuration changes, resource allocation adjustments, log file manipulations, and network traffic modifications;

defining at least one objective of the at least one agent to be performed, to the at least one simulated IT environment, by the agent device of the at least one agent and the at least one action option to be simulated for execution by the agent device of the at least one agent to perform the at least one objective;

repeatedly simulating, using a simulation algorithm executed by the processor, execution of the at least one action by the agent device of the at least one agent to the at least one simulated IT environment, and obtaining at least one simulated result of the simulated execution of the at least one action;

after each simulation of the execution of at least one action, detecting a deviation between the at least one simulated result of the simulated execution of the at least one action to the at least one simulated IT environment and the at least one objective of the at least one agent, and dynamically reconfiguring, via the processor, in response to the detected deviation, the at least one action determination model by replacing a first one of the plurality of ML-based AI neural net models having a first set of weighted parameters and stored in the storage device with a second one of the plurality of ML-based AI neural net models being different and separately stored from the first one of the plurality of ML-based AI neural net models and having a second set of weighted parameters different from the first set of weighted parameters, until the simulated result satisfies the at least one objective based on at least one of system performance metrics, network and software outputs and responses, error rates, and execution logs, thereby reducing the deviation in subsequent simulations,

wherein replacing the first one of the plurality of ML-based AI neural net models with the second one of the plurality of ML-based AI neural net models occurs as part of an interactive deviation-reduction loop.

2 . The method according to claim 1 , wherein dynamically reconfiguring, via the processor, the at least one action determination model involves optimizers, machine learning, or reinforcement learning.

3 . The method according to claim 1 , wherein the at least one action determination model defines at least one subsequent action option to be simulated for execution by the agent device of the at least one agent after the at least one action option currently simulated by the agent device of the at least one agent, to select in the at least one given context of attributes, state and/or behaviors of the at least one simulated IT environment.

4 . The method according to claim 1 , wherein the at least one action determination model determines at least one of an attacker action, an attack action, an exploit action, a defender action, a defending action, a detection action, a mitigation action, a prevention action, an alarm/alert action, a monitoring action, an evaluator action, a tester action, a penetration testing action, a vulnerability assessment action, a recommendation for human users, a configuration action for machines, a policy based action, a rule-based action, a user input action, a user output action, a data ingestion action, a repair action, assembly/disassembly action, a preparing action, a use action, a disposal action, a maintenance action, a directing action, an informational action, an entertaining action, a diagnosing action, transaction action, a purchasing action, a selling action, decision action, training action, education action, buying action, notification action, deception action, distraction action, timing action, delay action, support action, redirect action, or a transfer action.

5 . The method according to claim 1 , wherein the at least one simulated result includes at least one of an environment effect, an agent effect, console output, data returned by the action, data returned by the at least one simulated IT environment, context data, metadata, or data returned by the agent.

6 . The method according to claim 1 , wherein the at least one objective of the at least one agent includes at least one of successful attacking, defending, preventing, assessing, testing, evaluating, alarming, monitoring of the at least one simulated IT environment, providing a service, maintaining, updating, analyzing, deceiving, action execution, or action sequence execution.

7 . The method according to claim 1 , wherein the IT-system operation includes at least one of inputting, outputting, processing, transmitting, receiving and/or configuring data associated with or used for operation of the IT systems under the at least one simulated IT environment.

8 . The method according to claim 1 , wherein the IT-system operation is configured to perform network scanning, network configuration, policy enforcement, vulnerability assessment, and/or monitoring network status in the at least one simulated IT environment.