Systems and methods for controlling pallets in a manufacturing environment using reinforcement learning
A method includes obtaining sensor data from a plurality of sensors disposed at a plurality of routing control locations of an environment, where the sensor data is indicative of a number of a plurality of pallets at the plurality of routing control locations. The method includes calculating a plurality of difference values based on the sensor data, calculating a transient production value based on the sensor data and a transient objective function, and calculating a steady state production value based on the sensor data and a steady state objective function. The method includes generating a state vector based on the plurality of difference values, the transient production value, and the steady state production value, and defining a set of routes for a set of pallets from among the plurality of pallets based on the state vector and a digital twin of the environment.
1 . A method comprising:
obtaining sensor data from a plurality of sensors disposed at a plurality of routing control locations of an environment, wherein the sensor data is indicative of a number of a plurality of pallets at the plurality of routing control locations;
calculating a plurality of difference values based on the sensor data;
calculating a transient production value based on the sensor data and a transient objective function;
calculating a steady state production value based on the sensor data and a steady state objective function;
generating, using a reinforcement learning system including a neural network, wherein the neural network comprises a dueling network architecture that separately learns state values and action advantages, a state vector by combining the plurality of difference values, the transient production value, and the steady state production value into a unified vector representation that is used by the reinforcement learning system to select actions that optimize both transient and steady-state manufacturing objectives simultaneously based on the plurality of difference values, the transient production value, and the steady state production value;
defining, using the reinforcement learning system, a set of routes for a set of pallets from among the plurality of pallets based on the state vector and a digital twin of the environment, wherein the digital twin performs discrete event simulation synchronized with real-time sensor data to generate virtual event traces for training, wherein the digital twin of the environment includes a plurality of trained routes that are trained by the reinforcement learning system; and
controlling, using the reinforcement learning system, an autonomous movement of the set of pallets based on a real-time comparison between dynamically generated routes and the plurality of trained routes to prevent bottlenecks in the environment.
2 . The method of claim 1 , wherein the plurality of difference values corresponds to a difference between a first number of the plurality of pallets at a first routing control location from among the plurality of routing control locations and a second number of the plurality of pallets at a second routing control location from among the plurality of routing control locations.
3 . The method of claim 1 further comprising determining an action at each routing control location from among the plurality of routing control locations, wherein the action includes one of a pallet merging operation and a pallet splitting operation, and wherein the set of routes is further based on the action.
4 . The method of claim 1 , wherein the method further comprises:
identifying a set of trained routes from among the plurality of trained routes based on the state vector; and
defining the set of routes based on the set of trained routes.
5 . The method of claim 1 , wherein the method further comprises:
determining whether the set of routes correspond to a set of trained routes from among the plurality of trained routes; and
performing a corrective action in response to the set of routes not corresponding to the set of trained routes.
6 . The method of claim 5 , wherein performing the corrective action includes broadcasting a notification indicating that a manufacturing routine that utilizes the set of routes for pallet movement does not satisfy a time criterion.
7 . The method of claim 5 , wherein the corrective action includes performing a reinforcement training routine configured to selectively adjust one or more reinforcement parameters of the digital twin.
8 . The method of claim 7 , wherein selectively adjusting the one or more reinforcement parameters includes adjusting a sensor layout of a plurality of virtual sensors corresponding to the plurality of sensors, adjusting the transient objective function, adjusting the steady state objective function, or a combination thereof.
9 . The method of claim 1 , wherein the transient objective function and the steady state objective function are based on a first target production value for a first type of pallets from among the plurality of pallets and a second target production value for a second type of pallets from among the plurality of pallets.
10 . The method of claim 9 , wherein the transient objective function is based on a first number of the first type of pallets and a second number of the second type of pallets located between a first routing control location and a consequent routing control location.
11 . The method of claim 9 , wherein the steady state objective function is based on a first number of the first type of pallets and a second number of the second type of pallets located between a consequent routing subsequent routing control location downstream from the first routing control location and a destination.
12 . A system comprising:
a processor; and
a nontransitory computer-readable medium including instructions that are executable by the processor, wherein the instructions include:
obtaining sensor data from a plurality of sensors disposed at a plurality of routing control locations of an environment, wherein the sensor data is indicative of a number of a plurality of pallets at the plurality of routing control locations;
calculating a plurality of difference values based on the sensor data;
calculating a transient production value based on the sensor data and a transient objective function;
calculating a steady state production value based on the sensor data and a steady state objective function, wherein the transient objective function and the steady state objective function are based on a first target production value and a second target production value, wherein the first target production value is a first number of a first type of pallets from among the plurality of pallets that traverse between two locations of the environment to satisfy a constraint, and the second target production value is a second number of a second type of pallets from among the plurality of pallets that traverse between the two locations to satisfy the constraint;
generating a state vector by combining the plurality of difference values, the transient production value, and the steady state production value into a unified vector representation that is used by a reinforcement learning system to select actions that optimize both transient and steady-state manufacturing objectives simultaneously based on the plurality of difference values, the transient production value, and the steady state production value;
determining an action at each routing control location from among the plurality of routing control locations based on the state vector, wherein the action includes one of a pallet merging operation and a pallet splitting operation;
defining, using the reinforcement learning system including a neural network, wherein the neural network comprises a dueling network architecture that separately learns state values and action advantages, a set of routes for a set of pallets from among the plurality of pallets based on the action at each routing control location and a digital twin of the environment, wherein the digital twin performs discrete event simulation synchronized with real-time sensor data to generate virtual event traces for training, wherein the digital twin of the environment includes a plurality of trained routes for controlling movement of the plurality of pallets;
determining, using the reinforcement learning system, whether the set of routes correspond to a set of trained routes from among the plurality of trained routes;
performing, using the reinforcement learning system, a corrective action in response to the set of routes not corresponding to the set of trained routes, wherein the corrective action includes updating the set of routes to generate a new set of trained routes; and
controlling, using the reinforcement learning system, an autonomous movement of the set of pallets based on a real-time comparison between the new set of routes and the plurality of trained routes to prevent bottlenecks in the environment.
13 . The system of claim 12 , wherein the plurality of difference values corresponds to a difference between a third number of the plurality of pallets at a first routing control location from among the plurality of routing control locations and a fourth number of the plurality of pallets at a second routing control location from among the plurality of routing control locations.
14 . The system of claim 12 , wherein performing the corrective action includes broadcasting a notification indicating that a manufacturing routine that utilizes the set of routes for pallet movement does not satisfy a time criterion.
15 . The system of claim 12 , wherein the corrective action includes performing a reinforcement training routine configured to selectively adjust one or more reinforcement parameters of the digital twin.
16 . The system of claim 15 , wherein selectively adjusting the one or more reinforcement parameters includes adjusting a sensor layout of a plurality of virtual sensors corresponding to the plurality of sensors, adjusting the transient objective function, adjusting the steady state objective function, or a combination thereof.
17 . A method comprising:
obtaining sensor data from a plurality of sensors disposed at a plurality of routing control locations of an environment, wherein the sensor data is indicative of a number of a plurality of pallets at the plurality of routing control locations;
calculating a plurality of difference values based on the sensor data;
calculating a transient production value based on the sensor data and a transient objective function;
calculating a steady state production value based on the sensor data and a steady state objective function, wherein the transient objective function and the steady state objective function are based on a first target production value and a second target production value, wherein the first target production value is a first number of a first type of pallets from among the plurality of pallets that traverse between two locations of the environment to satisfy a constraint, and the second target production value is a second number of a second type of pallets from among the plurality of pallets that traverse between the two locations to satisfy the constraint;
generating a state vector by combining the plurality of difference values, the transient production value, and the steady state production value into a unified vector representation that is used by a reinforcement learning system to select actions that optimize both transient and steady-state manufacturing objectives simultaneously based on the plurality of difference values, the transient production value, and the steady state production value;
determining an action at each routing control location from among the plurality of routing control locations based on the state vector, wherein the action includes one of a pallet merging operation and a pallet splitting operation;
defining, using the reinforcement learning system including a neural network, wherein the neural network comprises a dueling network architecture that separately learns state values and action advantages, a set of routes for a set of pallets from among the plurality of pallets based on the action at each routing control location and a digital twin of the environment, wherein the digital twin performs discrete event simulation synchronized with real-time sensor data to generate virtual event traces for training, wherein the digital twin of the environment includes a plurality of trained routes for controlling movement of the plurality of pallets;
determining, using the reinforcement learning system, whether the set of routes correspond to a set of trained routes from among the plurality of trained routes;
performing, using the reinforcement learning system, a corrective action in response to the set of routes not corresponding to the set of trained routes, wherein the corrective action includes updating the set of routes to generate a new set of trained routes; and
controlling, using the reinforcement learning system, an autonomous movement of the set of pallets based on a real-time comparison between the new set of routes and the plurality of trained routes to prevent bottlenecks in the environment.
18 . The method of claim 17 , wherein performing the corrective action includes broadcasting a notification indicating that a manufacturing routine corresponding to the set of routes does not satisfy a time criteria.
19 . The method of claim 17 , wherein performing the corrective action includes broadcasting a notification indicating that a manufacturing routine that utilizes the set of routes for pallet movement does not satisfy a time criterion.
20 . The method of claim 17 , wherein the plurality of difference values correspond to a difference between a third number of the plurality of pallets at a first routing control location from among the plurality of routing control locations and a fourth number of the plurality of pallets at a second routing control location from among the plurality of routing control locations.