Device, computer program and computer-implemented method for machine learning
A device, computer program and computer-implemented method for machine learning. The method comprises providing a task comprising an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over rewards that is conditioned on actions, providing a hyperprior, wherein the hyperprior is a distribution over the action space, determining, depending on the hyperprior, a hyperposterior for that a lower bound for an expected reward on future bandit tasks has as large a value as possible, when using priors sampled from the hyperposterior, and wherein the hyperposterior is a distribution over the action space.
1 . A computer-implemented method for machine learning, the method comprising the following steps:
providing a task including an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over rewards that is conditioned on actions;
providing a hyperprior, wherein the hyperprior is a distribution over the action space;
initializing a posterior with a prior;
determining from a set of behavior policies of the task a behavior policy that is associated with the task posterior, wherein the behavior policy includes a distribution over actions with a probability mass;
sampling or selecting randomly from the probability mass an action;
sampling a reward depending on the action from the distribution over rewards;
determining a task dataset that includes the action and the reward;
updating the posterior to include the determined task dataset;
determining, depending on the hyperprior, a hyperposterior so that a lower bound for an expected reward on future bandit tasks is maximized, when using priors sampled from the hyperposterior, wherein the hyperposterior is a distribution over the action space;
processing sensor data including digital image data or audio data, depending on a prior sampled from the hyperposterior, for classifying the sensor data; and
(i) detecting presence of objects in the sensor data, or (ii) performing a semantic segmentation on the sensor data, or (iii) determining a measure for robustness of the machine learning including a probability that an expected error on a next task is not above a predetermined value, when sampling priors from the hyperposterior, or (iv) detecting an anomaly in sensor data depending on a prior sampled from the hyperposterior, or (v) learning a policy for controlling a physical system and determining a control signal for controlling the physical system depending on a prior sampled from the hyperposterior.
2 . The method according to claim 1 , further comprising:
determining the hyperposterior in a number of iterations; and
sampling a prior of an iteration from the hyperposterior of a previous iteration.
3 . The method according to claim 2 , further comprising:
sampling the task of the iteration from a distribution over tasks.
4 . The method according to claim 1 , further comprising:
initializing the hyperposterior with the hyperprior.
5 . The method according to claim 1 , wherein the providing the task includes providing the task including a state space and a distribution over initial states, and wherein the method further comprises sampling or selecting randomly an initial state from the distribution over initial states, and wherein the distribution over rewards is conditioned on the actions and states of the state space.
6 . The method according to claim 1 , the method further comprising:
initializing the task dataset with an empty set and then updating the posterior in a predetermined number of rounds.
7 . The method according to claim 1 , wherein the determining of the hyperposterior includes determining an approximation of an expected reward depending on a Kullback-Leibler divergence of the hyperposterior and the hyperprior.
8 . A device for machine learning, the device configured to:
provide a task including an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over rewards that is conditioned on actions;
provide a hyperprior, wherein the hyperprior is a distribution over the action space;
initialize a posterior with a prior;
determine from a set of behavior policies of the task a behavior policy that is associated with the task posterior, wherein the behavior policy includes a distribution over actions with a probability mass;
sample or select randomly from the probability mass an action;
sample a reward depending on the action from the distribution over rewards;
determine a task dataset that includes the action and the reward;
update the posterior to include the determined task dataset;
determine, depending on the hyperprior, a hyperposterior so that a lower bound for an expected reward on future bandit tasks is maximized, when using priors sampled from the hyperposterior, wherein the hyperposterior is a distribution over the action space;
process sensor data including digital image data or audio data, depending on a prior sampled from the hyperposterior, for classifying the sensor data; and
(i) detect presence of objects in the sensor data, or (ii) perform a semantic segmentation on the sensor data, or (iii) determine a measure for robustness of the machine learning including a probability that an expected error on a next task is not above a predetermined value, when sampling priors from the hyperposterior, or (iv) detect an anomaly in sensor data depending on a prior sampled from the hyperposterior, or (v) learn a policy for controlling a physical system and determining a control signal for controlling the physical system depending on a prior sampled from the hyperposterior.
9 . A non-transitory computer-readable medium on which is stored a computer program including computer-readable instructions for machine learning, the instructions, when executed by a computer, causing the computer to perform the following steps:
providing a task including an action space of a multi-armed bandit problem or a contextual bandit problem and a distribution over rewards that is conditioned on actions;
providing a hyperprior, wherein the hyperprior is a distribution over the action space;
initializing a posterior with a prior;
determining from a set of behavior policies of the task a behavior policy that is associated with the task posterior, wherein the behavior policy includes a distribution over actions with a probability mass;
sampling or selecting randomly from the probability mass an action;
sampling a reward depending on the action from the distribution over rewards;
determining a task dataset that includes the action and the reward;
updating the posterior to include the determined task dataset;
determining, depending on the hyperprior, a hyperposterior so that a lower bound for an expected reward on future bandit tasks is maximized, when using priors sampled from the hyperposterior, wherein the hyperposterior is a distribution over the action space;
processing sensor data including digital image data or audio data, depending on a prior sampled from the hyperposterior, for classifying the sensor data; and
(i) detecting presence of objects in the sensor data, or (ii) performing a semantic segmentation on the sensor data, or (iii) determining a measure for robustness of the machine learning including a probability that an expected error on a next task is not above a predetermined value, when sampling priors from the hyperposterior, or (iv) detecting an anomaly in sensor data depending on a prior sampled from the hyperposterior, or (v) learning a policy for controlling a physical system and determining a control signal for controlling the physical system depending on a prior sampled from the hyperposterior.