Learning constraints over beliefs in autonomous vehicle operations
Autonomous vehicle operations use learned constraints over beliefs to traverse a vehicle transportation network. Sensor data and user demonstration data are used to determine a belief path using a partially observable Markov decision process (POMDP) model. The belief path can be updated based on the learned constraints. Candidate actions are determined based on the POMDP model. The candidate actions are constrained by the updated belief path. An action is selected, and the vehicle traverses the vehicle network using the selected action.
1. A method for use in a vehicle, the method comprising:
obtaining sensor data;
obtaining user demonstration data, wherein the user demonstration data is data associated with a sequence of actions and probability distributions of a current state of the vehicle;
determining a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;
determining learned constraints by sampling a constraint and determining counter-factual belief path labels;
updating the belief path based on the learned constraints;
determining candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;
selecting an action of the candidate actions that is above a probability threshold; and
controlling the vehicle using the selected action to traverse a vehicle network.
2. The method of claim 1 , wherein determining the learned constraints comprises:
determining the counter-factual belief path labels to increase an expectation value of the user demonstration data;
selecting a subset of the determined counter-factual belief path labels; and
determining a candidate constraint based on the subset of the determined counter-factual belief path labels.
3. The method of claim 1 , wherein the POMDP model is based on a belief simplex.
4. The method of claim 3 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.
5. The method of claim 1 , wherein the updated belief path is a probability distribution over a vehicle state.
6. The method of claim 1 , further comprising:
determining whether the constraint is true for a belief path and action pair; and
disallowing the action at the belief path of the belief path and action pair when it is determined that the constraint is true.
7. The method of claim 1 , wherein the selected action is a stop action, and edge action, or a go action.
8. A vehicle, comprising:
a sensor configured to obtain sensor data;
a processor configured to:
obtain user demonstration data, wherein the user demonstration data is data associated with a sequence of actions and probability distributions of a current state of the vehicle, wherein the current state of the vehicle is determined based on the sensor data;
determine a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;
sample a constraint and determine counter-factual belief path labels to determine learned constraints;
update the belief path based on the learned constraints;
determine candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;
select an action of the candidate actions that is above a probability threshold; and
control the vehicle using the selected action to traverse a vehicle network.
9. The vehicle of claim 8 , wherein the processor is further configured to:
determine the counter-factual believe path labels to increase an expectation value of the user demonstration data;
select a subset of determined counter-factual belief path labels; and
determine a candidate constraint based on the subset of the determined counter-factual belief path labels.
10. The vehicle of claim 8 , wherein the POMDP model is based on a belief simplex.
11. The vehicle of claim 10 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.
12. The vehicle of claim 8 , wherein the updated belief path is a probability distribution over a vehicle state.
13. The vehicle of claim 8 , wherein the processor is further configured to:
determine whether the constraint is true for a belief path and action pair; and
disallow the action at the belief path of the belief path and action pair when it is determined that the constraint is true.
14. The vehicle of claim 8 , wherein the selected action is a stop action, an edge action, or a go action.
15. A non-transitory computer-readable medium comprising instructions, that when executed by a processor, cause the processor to perform operations comprising:
obtaining sensor data;
obtaining user demonstration data, wherein the user demonstration data is associated with a sequence of actions and probability distributions of a current state of a vehicle;
determining a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;
determining learned constraints by sampling a constraint and determining counter-factual belief path labels;
updating the belief path based on the learned constraints;
determining candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;
selecting an action of the candidate actions that is above a probability threshold; and
controlling the vehicle using the selected action to traverse a vehicle network.
16. The non-transitory computer-readable medium of claim 15 , the operations further comprising:
determining the counter-factual belief path labels to increase an expectation value of the user demonstration data;
selecting a subset of the determined counter-factual belief path labels; and
determining a candidate constraint based on the subset of the determined counter-factual belief path labels.
17. The non-transitory computer-readable medium of claim 15 , wherein the POMDP model is based on a belief simplex.
18. The non-transitory computer-readable medium of claim 17 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.
19. The non-transitory computer-readable medium of claim 15 , wherein the updated belief path is a probability distribution over a vehicle state.
20. The non-transitory computer-readable medium of claim 15 , the operations further comprising:
determining whether the constraint is true for a belief path and action pair; and
disallowing the action at the belief path of the belief path and action pair when it is determined that the constraint is true.