Computing robust policies in offline reinforcement learning
According to one embodiment, a method, computer system, and computer program product for reinforcement learning is provided. The present invention may include training, using an offline dataset, a plurality of diverse reward models, and creating a policy based on an output of the reward models and a robustness operator of the reward models.
1 . A computer implemented method for reinforcement learning, the method comprising:
generating, by a processor, a set of diverse reward models, wherein the diverse reward models comprise a plurality of reinforcement learning models differing from each other in at least one of hyper-parameters and machine learning technique;
training, by the processor, the diverse reward models in parallel to predict expected rewards for performing any of a plurality of state transitions and a plurality of actions using an offline dataset, wherein the offline dataset comprises the actions and the state transitions as performed by an agent and associated rewards;
determining, by the processor using the plurality of diverse reward models, a robustness operator, wherein the robustness operator is a function which takes as its input a vector of the expected rewards predicted by the set of diverse reward models for any of the actions and the state transitions, and expresses the vector as a single number;
computing, by the processor, a value function based on the robustness operator that determines how much reward the agent receives for the state transitions and the actions; and
creating, by the processor, a policy by utilizing the value function to determine a sequence of the state transitions and the actions resulting in a maximum total reward, wherein the policy expresses the sequence.
2 . The method of claim 1 , further comprising:
visualizing a behavior of the policy to a user.
3 . The method of claim 1 , further comprising:
adjusting one or more of the hyper-parameters associated with the diverse reward models based on user feedback.
4 . The method of claim 1 , wherein the plurality of diverse reward models are selected from regression models and neural networks.
5 . The method of claim 1 , wherein the robustness operator comprises a tau-percentile over a k-dimensional vector of reward values.
6 . The method of claim 1 , wherein the robustness operator comprises a minimum over a k-dimensional vector of reward values.
7 . The method of claim 1 , wherein the robustness operator comprises a preference relation defined by a convex combination of elements in a k-dimensional reward vector.
8 . A computer system for reinforcement learning, the computer system comprising:
one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the program instructions, when executed by the one or more processors, cause the computer system to perform a method comprising:
generating a set of diverse reward models, wherein the diverse reward models comprise a plurality of reinforcement learning models differing from each other in at least one of hyper-parameters and machine learning technique;
training the diverse reward models in parallel to predict expected rewards for performing any of a plurality of state transitions and a plurality of actions using an offline dataset, wherein the offline dataset comprises the actions and the state transitions as performed by an agent and associated rewards;
determining, using the plurality of diverse reward models, a robustness operator, wherein the robustness operator is a function which takes as its input a vector of the expected rewards predicted by the set of diverse reward models for any of the actions and the state transitions, and expresses the vector as a single number;
computing a value function based on the robustness operator that determines how much reward the agent receives for the state transitions and the actions; and
creating a policy by utilizing the value function to determine a sequence of the state transitions and the actions resulting in a maximum total reward, wherein the policy expresses the sequence.
9 . The computer system of claim 8 , further comprising:
visualizing a behavior of the policy to a user.
10 . The computer system of claim 8 , further comprising:
adjusting one or more of the hyper-parameters associated with the diverse reward models based on user feedback.
11 . The computer system of claim 8 , wherein the plurality of diverse reward models are selected from regression models and neural networks.
12 . The computer system of claim 8 , wherein the robustness operator comprises a tau-percentile over a k-dimensional vector of reward values.
13 . The computer system of claim 8 , wherein the robustness operator comprises a minimum over a k-dimensional vector of reward values.
14 . The computer system of claim 8 , wherein the robustness operator comprises a preference relation defined by a convex combination of elements in a k-dimensional reward vector.
15 . A computer program product for reinforcement learning, the computer program product comprising:
one or more computer-readable tangible storage medium and program instructions stored on at least one of the one or more tangible storage medium, the program instructions executable by a processor to cause the processor to perform a method comprising:
generating a set of diverse reward models, wherein the diverse reward models comprise a plurality of reinforcement learning models differing from each other in at least one of hyper-parameters and machine learning technique;
training the diverse reward models in parallel to predict expected rewards for performing any of a plurality of state transitions and a plurality of actions using an offline dataset, wherein the offline dataset comprises the actions and the state transitions as performed by an agent and associated rewards;
determining, using the plurality of diverse reward models, a robustness operator, wherein the robustness operator is a function which takes as its input a vector of the expected rewards predicted by the set of diverse reward models for any of the actions and the state transitions, and expresses the vector as a single number;
computing a value function based on the robustness operator that determines how much reward the agent receives for the state transitions and the actions; and
creating a policy by utilizing the value function to determine a sequence of the state transitions and the actions resulting in a maximum total reward, wherein the policy expresses the sequence.
16 . The computer program product of claim 15 , further comprising:
visualizing a behavior of the policy to a user.
17 . The computer program product of claim 15 , further comprising:
adjusting one or more of the hyper-parameters associated with the diverse reward models based on user feedback.
18 . The computer program product of claim 15 , wherein the plurality of diverse reward models are selected from regression models and neural networks.
19 . The computer program product of claim 15 , wherein the robustness operator comprises a tau-percentile over a k-dimensional vector of reward values.
20 . The computer program product of claim 15 , wherein the robustness operator comprises a minimum over a k-dimensional vector of reward values.