IP Library Granted Patent US 12,391,274
Granted Patent B2
US 12,391,274 · App. 18/175,747 · Granted Aug 19, 2025

Learning constraints over beliefs in autonomous vehicle operations

Inventors: Marcell Vazquez-Chanlatte (Palo Alto, CA); Stefan Witwicki (San Carlos, CA)
Assignee: Nissan North America, Inc.
B60W60/001B60W50/00G05B13/0265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,391,274
App. No.
18/175,747
Granted
Aug 19, 2025
Kind
B2
Abstract

Autonomous vehicle operations use learned constraints over beliefs to traverse a vehicle transportation network. Sensor data and user demonstration data are used to determine a belief path using a partially observable Markov decision process (POMDP) model. The belief path can be updated based on the learned constraints. Candidate actions are determined based on the POMDP model. The candidate actions are constrained by the updated belief path. An action is selected, and the vehicle traverses the vehicle network using the selected action.

Claims (60)

1. A method for use in a vehicle, the method comprising:

obtaining sensor data;

obtaining user demonstration data, wherein the user demonstration data is data associated with a sequence of actions and probability distributions of a current state of the vehicle;

determining a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;

determining learned constraints by sampling a constraint and determining counter-factual belief path labels;

updating the belief path based on the learned constraints;

determining candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;

selecting an action of the candidate actions that is above a probability threshold; and

controlling the vehicle using the selected action to traverse a vehicle network.

2. The method of claim 1 , wherein determining the learned constraints comprises:

determining the counter-factual belief path labels to increase an expectation value of the user demonstration data;

selecting a subset of the determined counter-factual belief path labels; and

determining a candidate constraint based on the subset of the determined counter-factual belief path labels.

3. The method of claim 1 , wherein the POMDP model is based on a belief simplex.

4. The method of claim 3 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.

5. The method of claim 1 , wherein the updated belief path is a probability distribution over a vehicle state.

6. The method of claim 1 , further comprising:

determining whether the constraint is true for a belief path and action pair; and

disallowing the action at the belief path of the belief path and action pair when it is determined that the constraint is true.

7. The method of claim 1 , wherein the selected action is a stop action, and edge action, or a go action.

8. A vehicle, comprising:

a sensor configured to obtain sensor data;

a processor configured to:

obtain user demonstration data, wherein the user demonstration data is data associated with a sequence of actions and probability distributions of a current state of the vehicle, wherein the current state of the vehicle is determined based on the sensor data;

determine a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;

sample a constraint and determine counter-factual belief path labels to determine learned constraints;

update the belief path based on the learned constraints;

determine candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;

select an action of the candidate actions that is above a probability threshold; and

control the vehicle using the selected action to traverse a vehicle network.

9. The vehicle of claim 8 , wherein the processor is further configured to:

determine the counter-factual believe path labels to increase an expectation value of the user demonstration data;

select a subset of determined counter-factual belief path labels; and

determine a candidate constraint based on the subset of the determined counter-factual belief path labels.

10. The vehicle of claim 8 , wherein the POMDP model is based on a belief simplex.

11. The vehicle of claim 10 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.

12. The vehicle of claim 8 , wherein the updated belief path is a probability distribution over a vehicle state.

13. The vehicle of claim 8 , wherein the processor is further configured to:

determine whether the constraint is true for a belief path and action pair; and

disallow the action at the belief path of the belief path and action pair when it is determined that the constraint is true.

14. The vehicle of claim 8 , wherein the selected action is a stop action, an edge action, or a go action.

15. A non-transitory computer-readable medium comprising instructions, that when executed by a processor, cause the processor to perform operations comprising:

obtaining sensor data;

obtaining user demonstration data, wherein the user demonstration data is associated with a sequence of actions and probability distributions of a current state of a vehicle;

determining a belief path based on the sensor data and the user demonstration data using a partially observable Markov decision process (POMDP) model;

determining learned constraints by sampling a constraint and determining counter-factual belief path labels;

updating the belief path based on the learned constraints;

determining candidate actions based on the POMDP model, wherein the candidate actions are constrained by the updated belief path;

selecting an action of the candidate actions that is above a probability threshold; and

controlling the vehicle using the selected action to traverse a vehicle network.

16. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

determining the counter-factual belief path labels to increase an expectation value of the user demonstration data;

selecting a subset of the determined counter-factual belief path labels; and

determining a candidate constraint based on the subset of the determined counter-factual belief path labels.

17. The non-transitory computer-readable medium of claim 15 , wherein the POMDP model is based on a belief simplex.

18. The non-transitory computer-readable medium of claim 17 , wherein the belief simplex is a geometric representation of probability distributions of the candidate actions over a finite set of outcomes.

19. The non-transitory computer-readable medium of claim 15 , wherein the updated belief path is a probability distribution over a vehicle state.

20. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

determining whether the constraint is true for a belief path and action pair; and

disallowing the action at the belief path of the belief path and action pair when it is determined that the constraint is true.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2026
From: NISSAN NORTH AMERICA, INC.
To: NISSAN MOTOR CO., LTD.
Reel/Frame 074681/0544 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2023
From: VAZQUEZ-CHANLATTE, MARCELL; WITWICKI, STEFAN
To: NISSAN NORTH AMERICA, INC.
Reel/Frame 063042/0743 →
Continuity (1)
Related Publication 20240286634A1 · Aug 29, 2024
References Cited (31)
US 10649453B1 · Svegliato · 2020 [cited by examiner]
US 11027751B2 · Wray et al. · 2021 [cited by applicant]
US 11332165B2 · Akash · 2022 [cited by examiner]
US 11635758B2 · Wray · 2023 [cited by examiner]
US 20060200333A1 · Dalal · 2006 [cited by examiner]
US 20170297576A1 · Halder · 2017 [cited by examiner]
US 20190329763A1 · Sierra Gonzalez · 2019 [cited by examiner]
US 20200005645A1 · Wray · 2020 [cited by examiner]
US 20200073382A1 · Noda · 2020 [cited by examiner]
US 20200269875A1 · Wray · 2020 [cited by examiner]
US 20200331491A1 · Wray · 2020 [cited by examiner]
US 20210009154A1 · Wray · 2021 [cited by examiner]
US 20210078602A1 · Wray · 2021 [cited by examiner]
US 20210157315A1 · Wray · 2021 [cited by examiner]
US 20210188297A1 · Wray · 2021 [cited by examiner]
US 20210200208A1 · Wray · 2021 [cited by examiner]
US 20210268653A1 · Tian · 2021 [cited by examiner]
US 20220097736A1 · Lin · 2022 [cited by examiner]
US 20220164636A1 · Fadaie · 2022 [cited by examiner]
US 20220324484A1 · Hruschka · 2022 [cited by examiner]
US 20220371612A1 · Wray · 2022 [cited by examiner]
US 20220382279A1 · Wray · 2022 [cited by examiner]
US 20230227031A1 · Kobashi · 2023 [cited by examiner]
US 20240140472A1 · Bill Clark · 2024 [cited by examiner]
US 20240149920A1 · Zheng · 2024 [cited by examiner]
US 20240166242A1 · Dai · 2024 [cited by examiner]
US 20250083702A1 · Drusinsky · 2025 [cited by examiner]
Kyle Hollins Wray et al., POMDPs for Safe Visibility Reasoning in Autonomous Vehicles, Mar. 2021, IEEE, pp. 191-195. [cited by examiner]
Marcell Vazquez-Chanlatte; Specifications from Demonstrations: Learning, Teaching, and Control; EECS Department University of California, Berkeley Technical Report No. UCB/EECS-2022-107, May 13, 2022; http://www2.eecs.b… [cited by applicant]
Jha, Susmit & Seshia, Sanjit. (2014). Are there good mistakes? A theoretical analysis of CEGIS. Electronic Proceedings in Theoretical Computer Science. 157. 10.4204/EPTCS.157.10. ; https:/arxiv.org/pdf/1407.5397.pdf. [cited by applicant]
Software Modeling and Verification Group; RWTH Aachen University, Aachen, Germany; STORM; https://www.stormchecker.org/; accessed Feb. 28, 2023. [cited by applicant]