IP Library › Granted Patent US 12,524,704
Granted Patent B2
US 12,524,704 · App. 17/477,713 · Granted Jan 13, 2026

Learning based modeling of emergent behaviour of complex system

Inventors: Prasenjit Das (Kolkata, IN); Souvik Barat (Pune, IN); Vinay Kulkarni (Pune, IN); Prashant Kulmar (Pune, IN); Kaustav Bhattacharya (Chennai, IN); Sankaranarayanan Viswanathan (Chennai, IN)
Assignee: Tata Consultancy Services Limited
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,704
App. No.
17/477,713
Filed
Sep 17, 2021
Granted
Jan 13, 2026
Kind
B2
Art Unit
2144
USPC
706/12
Abstract

The disclosure generally related to a learning-based modelling of an emergent behavior of a complex system. Existing decision-making at complex systems primarily relies on qualitative approaches, which often results in inaccurate outputs. The disclosed system includes a digital twin of the complex system and a digital twin of an environment of said complex system, and captures an interaction and dynamic behavior of agents of the digital twins. The agents of the digital twins are simulated and modelled using learning-based models such as RL and genetic algorithms that learns the behavior (i.e. actions and their outcomes) over a period of time. Hence, the agents (or actors) of the digital twins are dynamic in nature. The actor-based bottom up simulation approach is capable of producing sufficient insight for effective decision making prior to implementation.

Claims (44)

1 . A processor-implemented method comprising:

simulating, via one or more hardware processors, a first digital twin of a complex system and a second digital twin of an environment associated with the complex system, the first digital twin comprising a first set of digitally configured dynamic agents and the second digital twin comprising a second set of digitally configured dynamic agents, wherein the first digital twin and the second digital twin are digital replica of the complex system and the environment associated with the complex system, wherein each of the first set and the second set of digitally configured dynamic agents defined using one or more state variables, one or more characteristic variables and a set of actions, wherein the first set of digitally configured dynamic agents and the second set of digitally configured dynamic agents are capable of learning over a period of time by observing the set of actions in the complex system and the environment respectively, wherein the first digital twin and the second digital twin are constructed by identifying information from plurality of sources of information of the complex system, wherein the first digital twin and second digital twin are validated by subjecting the first digital twin and second digital twin to past events to simulate past behavior, wherein one or more characteristics of the first set and the second set of digitally configured dynamic agents change over a time by observing the set of actions resulting into good or bad state over the time, wherein the complex system comprises a system of systems, and the complex system is a reactive entity and exchanges messages and resources with the environment associated with the complex system, the complex system composes a number of interdependent subsystems or elements in a nonlinear way, wherein the digital twins capture and represents structure as well as behavior of the complex system and the environment, and the digital twins capture an interaction and dynamic behavior of the digitally configured dynamic agents;

receiving, via the one or more hardware processors, a trigger at one or more digitally configured dynamic agents from amongst the first set of digitally configured dynamic agents and the second set of digitally configured dynamic agents;

computing, via the one or more hardware processors, a current value of the one or more state variables and the one or more characteristic variables associated with the one or more digitally configured dynamic agents by accessing a system database;

predicting, via the one or more hardware processors, a difference between the current value and an expected value of the one or more state variables and the one or more characteristics variables, the current value computed using a multi criteria decision making technique, the expected value obtained from one or more goals associated with the complex system, the one or more goals prestored in a knowledge repository associated with the complex system, and wherein the one or more goals are indicative of decision-making in response to the trigger in the complex system;

defining, based on the difference between the current value and the expected value of the one or more state variables and the one or more characteristic variables, a decision function for the one or more digitally configured dynamic agents using the decision function, via the one or more hardware processors, wherein the decision function is realized using at least one of a Reinforcement learning (RL) and optimization technique, and wherein an observation is indicative of outcome of the decision function and ability to reach to a desired state in a future time, wherein RL decides the characteristic variables which should be changed and the change to be introduced in the characteristic variables of the agent;

simulating iteratively, via the one or more hardware processors, the first and the second set of digitally configured dynamic agents based on the decision function in a plurality of iterations until the difference between the current value and the expected value of the one or more state variables is determined to be within a predetermined threshold limit, wherein each of the first set and the second set of digitally configured dynamic agents comprises an RL agent capable of observing action taken by the decision function and consequence of the action during the iterative simulations; and

triggering a modification in the one or more characteristics variables of a digitally configured dynamic agent from amongst the one or more digitally configured dynamic agents based on a series of actions taken by the digitally configured dynamic agent over a predefined period of time and observing values in terms of the one or more state variables in cognizance of the one or more goals of the digitally configured dynamic agent and actual state variable after the predefined period of time, wherein the actions taken by each of the digitally configured dynamic agent over the predefined period of time leads to learning of that dynamic agent,

wherein an action from amongst the set of actions is defined using a tuple comprising an event, the trigger, a computation function, and a resistance value, and

wherein the computation function is a function of the one or more state variables and the one or more characteristics variables, wherein the RL is used to decide the characteristic variables to be changed and the change to be introduced in the characteristic variables of the agent,

wherein the action is triggered by the agent when outcome of computation function is greater than the threshold value, which is a function of characteristic variables, wherein each of the digitally configured dynamic agent have different threshold to trigger the action, and

wherein the computation function is calculated using Multicriteria decision-making (MCDM) technique,

wherein the resistance comprises a threshold, wherein the threshold is a function over the one or more characteristics variables, and

enabling decision makers in evaluating different decision alternatives on the digital twin of the complex system and enabling to identify optimal choices to be applied on the complex system in an automated manner.

2 . A system comprising:

a memory storing instructions;

one or more communication interfaces; and

one or more hardware processors coupled to the memory via the one or more communication interfaces, wherein the one or more hardware processors are configured by the instructions to:

simulate a first digital twin of a complex system and a second digital twin of an environment associated with the complex system, the first digital twin comprising a first set of digitally configured dynamic agents and the second digital twin comprising a second set of digitally configured dynamic agents, wherein the first digital twin and the second digital twin are digital replica of the complex system and the environment associated with the complex system, each of the first set and the second set of digitally configured dynamic agents defined using one or more state variables, one or more characteristic variables and a set of actions, wherein the first set of digitally configured dynamic agents and the second set of digitally configured dynamic agents are capable of learning over a period of time by observing the set of actions in the complex system and the environment respectively, wherein the first digital twin and the second digital twin are constructed by identifying information from plurality of sources of information of the complex system, wherein the first digital twin and second digital twin are validated by subjecting the first digital twin and second digital twin to past events to simulate past behavior, wherein one or more characteristics of the set of digitally configured dynamic agents change over a time by observing the set of actions resulting into good or bad state over the time, wherein the complex system comprises a system of systems, and the complex system is a reactive entity and exchanges messages and resources with the environment associated with the complex system, the complex system composes a number of interdependent subsystems or elements in a nonlinear way, wherein the digital twins capture and represents structure as well as behavior of the complex system and the environment, and the digital twins capture an interaction and dynamic behavior of the digitally configured dynamic agents;

receive a trigger at one or more digitally configured dynamic agents from amongst the first set digitally configured dynamic agents and the second set of digitally configured dynamic agents;

compute a current value of the one or more state variables and the one or more characteristic variables associated with the one or more digitally configured dynamic agents by accessing a system database;

predict a difference between the current value and an expected value of the one or more state variables and the one or more characteristics variables, the current value computed using multi criteria decision making technique, the expected value obtained from one or more goals associated with the complex system, the one or more goals prestored in a knowledge repository associated with the complex system, and wherein the one or more goals are indicative of decision-making in response to the trigger in the complex system;

define, based on the difference between the current value and the expected value of the one or more state variables and the one or more characteristic variables, a decision function for the one or more digitally configured dynamic agents using the decision function, wherein the decision function is realized using at least one of a Reinforcement learning (RL) and optimization technique, and wherein the observation is indicative of outcome of the decision function and ability to reach to a desired state in a future time, wherein RL decides the characteristic variables which should be changed and the change to be introduced in the characteristic variables of the agent; and

simulate iteratively the first and the second set of digitally configured dynamic agents based on the decision function in a plurality of iterations until the difference between the current value and the expected value of the one or more state variables is determined to be within a predetermined threshold limit, wherein each of the first set and the second set of digitally configured dynamic agents comprises an RL agent capable of observing action taken by the decision function and consequence of the action during the iterative simulations; and

trigger a modification in the one or more characteristics variables of a digitally configured dynamic agent from amongst the one or more digitally configured dynamic agents based on a series of actions taken by the digitally configured dynamic agent over a predefined period of time and observing values in terms of the one or more state variables in cognizance of the one or more goals of the digitally configured dynamic agent and actual state variable after the predefined period of time, wherein the actions taken by each of the digitally configured dynamic agent over the predefined period of time leads to learning of that dynamic agent, wherein an action from amongst the set of actions is defined using a tuple comprising an event, the trigger, a computation function, and a resistance value, and

wherein the computation function is a function of the one or more state variables and the one or more characteristics variables, wherein the RL is used to decide the characteristic variables to be changed and the change to be introduced in the characteristic variables of the agent,

wherein the action is triggered by the agent when outcome of computation function is greater than the threshold value, which is a function of characteristic variables, wherein each of the digitally configured dynamic agent have different threshold to trigger the action, and

wherein the computation function is calculated using Multicriteria decision-making (MCDM) technique, and

wherein the resistance comprises a threshold, wherein the threshold is a function over the one or more characteristics variables, and

enabling decision makers in evaluating different decision alternatives on the digital twin of the complex system and enabling to identify optimal choices to be applied on the complex system in an automated manner.

3 . One or more non-transitory machine readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:

simulating, via one or more hardware processors, a first digital twin of a complex system and a second digital twin of an environment associated with the complex system, the first digital twin comprising a first set of digitally configured dynamic agents and the second digital twin comprising a second set of digitally configured dynamic agents, wherein the first digital twin and the second digital twin are digital replica of the complex system and the environment associated with the complex system, each of the first set and the second set of digitally configured dynamic agents defined using one or more state variables, one or more characteristic variables and a set of actions, wherein the first set of digitally configured dynamic agents and the second set of digitally configured dynamic agents are capable of learning over a period of time by observing the set of actions in the complex system and the environment respectively, wherein the first digital twin and the second digital twin are constructed by identifying information from plurality of sources of information of the complex system, wherein the first digital twin and second digital twin are validated by subjecting the first digital twin and second digital twin to past events to simulate past behavior, wherein one or more characteristics of the first set and the second set of digitally configured dynamic agents change over a time by observing the set of actions resulting into good or bad state over the time, wherein the complex system comprises a system of systems, and the complex system is a reactive entity and exchanges messages and resources with the environment associated with the complex system, the complex system composes a number of interdependent subsystems or elements in a nonlinear way, wherein the digital twins capture and represents structure as well as behavior of the complex system and the environment, and the digital twins capture an interaction and dynamic behavior of the digitally configured dynamic agents;

receiving, via the one or more hardware processors, a trigger at one or more digitally configured dynamic agents from amongst the first set digitally configured dynamic agents and the second set of digitally configured dynamic agents;

computing, via the one or more hardware processors, a current value of the one or more state variables and the one or more characteristic variables associated with the one or more digitally configured dynamic agents by accessing a system database;

predicting, via the one or more hardware processors, a difference between the current value and an expected value of the one or more state variables and the one or more characteristics variables, the current value computed using a multi criteria decision making technique, the expected value obtained from one or more goals associated with the complex system, the one or more goals prestored in a knowledge repository associated with the complex system, and wherein the one or more goals are indicative of decision-making in response to the trigger in the complex system;

defining, based on the difference between the current value and the expected value of the one or more state variables and the one or more characteristic variables, a decision function for the one or more digitally configured dynamic agents using the decision function, via the one or more hardware processors, wherein the decision function is realized using at least one of a Reinforcement learning (RL) and optimization technique, and wherein the observation is indicative of outcome of the decision function and ability to reach to a desired state in a future time, wherein RL decides the characteristic variables which should be changed and the change to be introduced in the characteristic variables of the agent; and

simulating iteratively, via the one or more hardware processors, the first and the second set of digitally configured dynamic agents based on the decision function in a plurality of iterations until the difference between the current value and the expected value of the one or more state variables is determined to be within a predetermined threshold limit, wherein each of the first set and the second set of digitally configured dynamic agents comprises an RL agent capable of observing action taken by the decision function and consequence of the action during the iterative simulations; and

triggering a modification in the one or more characteristics variables of a digitally configured dynamic agent from amongst the one or more digitally configured dynamic agents based on a series of actions taken by the digitally configured dynamic agent over a predefined period of time and observing values in terms of the one or more state variables in cognizance of the one or more goals of the digitally configured dynamic agent and actual state variable after the predefined period of time, wherein the actions taken by each of the digitally configured dynamic agent over the predefined period of time leads to learning of that dynamic agent,

wherein an action from amongst the set of actions is defined using a tuple comprising an event, the trigger, a computation function, and a resistance value, and

wherein the computation function is a function of the one or more state variables and the one or more characteristics variables, wherein the RL is used to decide the characteristic variables to be changed and the change to be introduced in the characteristic variables of the agent,

wherein the action is triggered by the agent when outcome of computation function is greater than the threshold value, which is a function of characteristic variables, wherein each of the digitally configured dynamic agent have different threshold to trigger the action, and

wherein the computation function is calculated using Multicriteria decision-making (MCDM) technique,

wherein the resistance comprises a threshold, wherein the threshold is a function over the one or more characteristics variables, and

enabling decision makers in evaluating different decision alternatives on the digital twin of the complex system and enabling to identify optimal choices to be applied on the complex system in an automated manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2021
From: DAS, PRASENJIT; BARAT, SOUVIK; KULKARNI, VINAY; KUMAR, PRASHANT; BHATTACHARYA, KAUSTAV; VISWANATHAN, SANKARANARAYANAN
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 057511/0475 →
Priority Claims (1)
IN 202121008298 · Feb 26, 2021 · national
Continuity (1)
Related Publication 20220318676A1 · Oct 6, 2022
References Cited (33)
US 9792397B1 · Nagaraja · 2017 [cited by examiner]
US 10282512B2 · Bennett · 2019 [cited by examiner]
US 10592828B2 · Farooq · 2020 [cited by examiner]
US 20070150330A1 · McGoveran · 2007 [cited by applicant]
US 20180240043A1 · Majumdar · 2018 [cited by examiner]
US 20190122092A1 · Haines · 2019 [cited by examiner]
US 20190171438A1 · Franchitti · 2019 [cited by examiner]
US 20200380434A1 · Bhattacharya · 2020 [cited by examiner]
US 20210357555A1 · Liu · 2021 [cited by examiner]
CN 110488629A · 2019 [cited by examiner]
CN 111241752A · 2020 [cited by examiner]
CN 111797163A · 2020 [cited by examiner]
JP 2019508830A · 2019 [cited by examiner]
WO WO2020229904A1 · 2020 [cited by applicant]
Souvik Charkraborty, etc., “Machine Learning Based Digital Twin for Dynamical Systems with Multiple Time-Scales”, published Jun. 14, 2020 to arXiv, retrieved Oct. 17, 2024. (Year: 2020). [cited by examiner]
Qiang Liu, etc., “Digital twin-based designing of the configuration, motion, control, and optimization model of a flow-type smart manufacturing system”, published in Journal of Manufacturing Systems, vol. 58, part B, pp… [cited by examiner]
James Moyne, etc., “A Requirements Driven Digital Twin Framework: Specification and Opportunities”, published in IEEE Access, vol. 8, Jun. 5, 2020, retrieved Oct. 17, 2024. (Year: 2020). [cited by examiner]
Christian Stary, etc., “Digital Twin Generation: Re-Conceptualizing Agent Systems for Behavior-Centered Cyber-Physical System Development”, published in Sensors 2021, 21, 1096, Feb. 5, 2021, retrieved Oct. 17, 2024. (Ye… [cited by examiner]
Leonardo A. Espinosa Leal, etc., “Autonomous Industrial Management via Reinforcement Learning: Self-Learning Agents for Decision-Making—A Review”, published on Oct. 20, 2019 to arXiv, retrieved Apr. 23, 2025. (Year: 201… [cited by examiner]
Angira Sharma, etc., “Digital Twins: State of the Art Theory and Practice, Challenges, and Open Research Questions”, published on Dec. 4, 2020 to arXiv, retrieved Apr. 23, 2025. (Year: 2020). [cited by examiner]
M. Mazhar Rathore, etc., “The Role of AI, Machine Learning, and Big Data in Digital Twinning: A Systematic Literature Review, Challenges, and Opportunities”, published Feb. 22, 2021 to IEEE Access, retrieved Apr. 23, 20… [cited by examiner]
Constantin Cronrath, etc., “Enhancing Digital Twins through Reinforcement Learning”, published to IEEE Xplore on Sep. 19, 2019, retrieved Apr. 23, 2025. (Year: 2019). [cited by examiner]
Zai Muller-Zhang, etc., “Dynamic Process Planning using Digital Twins and Reinforcement Learning”, published via 2020 25th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA), Sep. 8-11,… [cited by examiner]
Florian Jaensch, etc., “Digital Twins of Manufacturing Systems as a Base for Machine Learning”, published via 2018 25th International Conference on Mechatronics and Machine Vision in Practice (M2VIP), Nov. 20-22, 2018, … [cited by examiner]
Florian Jaensch, etc., “Reinforcement Learning of Material Flow Control Logic using Hardware-in-the-Loop Simulation”, published via 2018 First International Conference on Artificial Intelligence for Industries (AI4I), S… [cited by examiner]
Zijie Ren, etc., “Strengthening Digital Twin Applications based on Machine Learning for Complex Equipment”, published via 2021 Design, Automation & Test in Europe Conference & Exhibition (Date), Feb. 1-5, 2021, retrieve… [cited by examiner]
Raju Kandaswamy, “Digital Twins and Reinforcement Learning using Unity ML-Agents”, published to YouTube on Oct. 1, 2020 at https://www.youtube.com/watch?v=Kr3dx8QD8a4, retrieved Apr. 23, 2025. (Year: 2020). [cited by examiner]
Igor Kiselev, etc., “An Adaptive Multi-agent System for Continuous Learning of Streaming Data”, published via 2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (vol. 2, 2008… [cited by examiner]
Bruno Maione, etc., “Evolutionary Learning Agents for Shop Floor Control”, published via 1999 7th IEEE International Conference on Emerging Technologies and Factory Automation. Proceedings ETFA '99 (Cat. No. 99TH8467), … [cited by examiner]
Zhenglei He, etc., “A Deep Reinforcement Learning Based Multi-Criteria Decision Support System for Textile Manufacturing Process Optimization”, published to arXiv as of Dec. 29, 2020, retrieved Jul. 14, 2025. (Year: 202… [cited by examiner]
Sindhu Padakandla, etc., A Survey of Reinforcement Learning Algorithms for Dynamically Varying Environments, published to arXiv as of May 19, 2020, retrieved Jul. 14, 2025. (Year: 2020). [cited by examiner]
Mathworks, “Train Reinforcement Learning Agents—Matlab & Simulink”, published on Apr. 1, 2019 to https://www.mathworks.com/help/reinforcement-learning/ug/train-reinforcement-learning-agents.html, retrieved Aug. 13, 2025… [cited by examiner]
Souvik Barat, “An actor based simulation driven digital twin for analyzing complex business systems,” Winter Simulation Conference (WSC), Dec. 2019, IEEE, https://eprints.mdx.ac.uk/27284/1/Con228%20-%20WinterSim-correct… [cited by applicant]