IP Library Granted Patent US 12,299,554
Granted Patent B2
US 12,299,554 · App. 18/653,211 · Granted May 13, 2025

Method and apparatus for constructing informative outcomes to guide multi-policy decision making

Inventors: Edwin Olson (Ann Arbor, MI); Dhanvin H. Mehta (Ann Arbor, MI); Gonzalo Ferrer (Ann Arbor, MI)
Assignee: The Regents of The University of Michigan
G06N3/02G06N3/008G06N3/084G06N7/01H04N1/00002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,554
App. No.
18/653,211
Granted
May 13, 2025
Kind
B2
Abstract

In Multi-Policy Decision-Making (MPDM), many computationally-expensive forward simulations are performed in order to predict the performance of a set of candidate policies. In risk-aware formulations of MPDM, only the worst outcomes affect the decision making process, and efficiently finding these influential outcomes becomes the core challenge. Recently, stochastic gradient optimization algorithms, using a heuristic function, were shown to be significantly superior to random sampling. In this disclosure, it was shown that accurate gradients can be computed-even through a complex forward simulation—using approaches similar to those in dep networks. The proposed approach finds influential outcomes more reliably, and is faster than earlier methods, allowing one to evaluate more policies while simultaneously eliminating the need to design an easily-differentiable heuristic function.

Claims (79)

1. A method, comprising:

operating a vehicle according to a first policy;

while operating the vehicle according to the first policy, evaluating a set of policy options, comprising:

detecting a set of objects in the vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a second policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

operating the vehicle according to the second policy;

wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a progress of the vehicle toward a predetermined goal and each of the set of policy options is further evaluated based on a quantified risk of executing the policy option.

2. The method of claim 1 , wherein the particular outcome category comprises a category associated with a high risk of collision between one or more of: the vehicle and at least one object of the set of objects; or a first object and a second object of the set of objects.

3. The method of claim 2 , wherein guiding the set of multiple potential outcomes to the particular outcome category comprises adjusting input data associated with the set of objects.

4. The method of claim 3 , wherein the input data comprises at least one of position or motion information associated with the set of objects.

5. The method of claim 3 , wherein the input data comprises a goal associated with the set of objects.

6. The method of claim 3 , wherein guiding the set of multiple potential outcomes further comprises applying a backpropagation process.

7. The method of claim 1 , wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a quantified inconvenience to the set of objects in response to executing the policy option, and producing the score based on the quantified inconvenience.

8. The method of claim 7 , wherein the quantified inconvenience is calculated based on a set of predicted distances between the vehicle and a closest object of the set of objects.

9. The method of claim 7 , wherein each of the set of policy options is further evaluated based on a quantified risk of executing the policy option.

10. A system, comprising:

a processing subsystem of a controlled vehicle configured to:

evaluate a set of policy options, comprising:

detecting a set of objects in the controlled vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

a control subsystem configured to implement the selected policy;

wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a quantified inconvenience to the set of objects in response to executing the policy option, and producing the score based on the quantified inconvenience.

11. The system of claim 10 , wherein the particular outcome category comprises a category associated with a high risk of collision between one or more of: the controlled vehicle and at least one object of the set of objects; or a first object and a second object of the set of objects.

12. The system of claim 11 , wherein guiding the set of multiple potential outcomes to the particular outcome category comprises adjusting input data associated with the set of objects.

13. The system of claim 12 , wherein the input data comprises at least one of position or motion information associated with the set of objects.

14. The system of claim 12 , wherein the input data comprises a goal associated with the set of objects.

15. The system of claim 10 , wherein the quantified inconvenience is calculated based on a set of predicted distances between the controlled vehicle and a closest object of the set of objects.

16. The system of claim 10 , wherein each of the set of policy options is further evaluated based on a quantified risk of executing the policy option.

17. The system of claim 10 , wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a progress of the vehicle toward a predetermined goal.

18. The system of claim 10 , wherein guiding the set of multiple potential outcomes includes applying a backpropagation process.

19. The system of claim 10 wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a progress of the vehicle toward a predetermined goal and each of the set of policy options is further evaluated based on a quantified risk of executing the policy option.

20. A method, comprising:

operating a vehicle according to a first policy;

while operating the vehicle according to the first policy, evaluating a set of policy options, comprising:

detecting a set of objects in the vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a second policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

operating the vehicle according to the second policy;

wherein the particular outcome category comprises a category associated with a high risk of collision between one or more of: the vehicle and at least one object of the set of objects; or a first object and a second object of the set of objects; and;

wherein guiding the set of multiple potential outcomes to the particular outcome category comprises adjusting input data associated with the set of objects and applying a backpropagation process.

21. A method, comprising:

operating a vehicle according to a first policy;

while operating the vehicle according to the first policy, evaluating a set of policy options, comprising:

detecting a set of objects in the vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a second policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

operating the vehicle according to the second policy;

wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a quantified inconvenience to the set of objects in response to executing the policy option, and producing the score based on the quantified inconvenience.

22. A system, comprising:

a processing subsystem of a controlled vehicle configured to:

evaluate a set of policy options, comprising:

detecting a set of objects in the controlled vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

a control subsystem configured to implement the selected policy;

wherein the particular outcome category comprises a category associated with a high risk of collision between one or more of: the vehicle and at least one object of the set of objects; or a first object and a second object of the set of objects; and

wherein guiding the set of multiple potential outcomes to the particular outcome category comprises adjusting input data associated with the set of objects and applying a backpropagation process.

23. A system, comprising:

a processing subsystem of a controlled vehicle configured to:

evaluate a set of policy options, comprising:

detecting a set of objects in the controlled vehicle's environment;

evaluating each of the set of policy options, comprising, for each of the set of policy options:

identifying a set of multiple potential outcomes associated with the policy, comprising guiding the set of multiple potential outcomes to a particular outcome category;

evaluating the set of multiple potential outcomes to produce a score;

selecting a policy from the set of policy options based on a set of scores comprising the produced score for each policy option; and

a control subsystem configured to implement the selected policy;

wherein evaluating the set of multiple potential outcomes comprises, for each of the set of policy options, predicting a progress of the vehicle toward a predetermined goal and each of the set of policy options is further evaluated based on a quantified risk of executing the policy option.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2025
From: OLSON, EDWIN; MEHTA, DHANVIN; FERRER, GONZALO
To: THE REGENTS OF THE UNIVERSITY OF MICHIGAN,
Reel/Frame 070348/0979 →
Continuity (5)
Continuation 18196897 · May 12, 2023
Continuation 17371221 · Jul 9, 2021
Continuation 15923577 · Mar 16, 2018
Provisional Application 62472734 · Mar 17, 2017
Related Publication 20240281635A1 · Aug 22, 2024
References Cited (91)
US 5544282A · Chen et al. · 1996 [cited by applicant]
US 6199013B1 · O'Shea · 2001 [cited by applicant]
US 8457827B1 · Ferguson · 2013 [cited by examiner]
US 8849557B1 · Levandowski · 2014 [cited by examiner]
US 9129519B2 · Aoude et al. · 2015 [cited by applicant]
US 9274525B1 · Ferguson et al. · 2016 [cited by applicant]
US 9495874B1 · Zhu et al. · 2016 [cited by applicant]
US 9618938B2 · Olson et al. · 2017 [cited by applicant]
US 9646428B1 · Konrardy et al. · 2017 [cited by applicant]
US 9720412B1 · Zhu et al. · 2017 [cited by applicant]
US 9811760B2 · Richardson et al. · 2017 [cited by applicant]
US 10012981B2 · Gariepy et al. · 2018 [cited by applicant]
US 10019005B2 · Wang · 2018 [cited by examiner]
US 10235882B1 · Aoude et al. · 2019 [cited by applicant]
US 10248120B1 · Siegel et al. · 2019 [cited by applicant]
US 10386856B2 · Wood et al. · 2019 [cited by applicant]
US 10540892B1 · Fields et al. · 2020 [cited by applicant]
US 10558224B1 · Lin et al. · 2020 [cited by applicant]
US 10586254B2 · Singhal · 2020 [cited by applicant]
US 10642276B2 · Huai · 2020 [cited by applicant]
US 10802484B2 · Jiang · 2020 [cited by examiner]
US 10969470B2 · Voorheis et al. · 2021 [cited by applicant]
US 11760387B2 · Raichelgauz · 2023 [cited by examiner]
US 11869092B2 · Konrardy et al. · 2024 [cited by applicant]
US 20020062207A1 · Faghri · 2002 [cited by applicant]
US 20040100563A1 · Sablak et al. · 2004 [cited by applicant]
US 20050004723A1 · Duggan et al. · 2005 [cited by applicant]
US 20060184275A1 · Hosokawa et al. · 2006 [cited by applicant]
US 20060200333A1 · Dalal et al. · 2006 [cited by applicant]
US 20070193798A1 · Allard et al. · 2007 [cited by applicant]
US 20070276600A1 · King et al. · 2007 [cited by applicant]
US 20080033684A1 · Vian et al. · 2008 [cited by applicant]
US 20100114554A1 · Misra · 2010 [cited by applicant]
US 20120089275A1 · Yao-Chang et al. · 2012 [cited by applicant]
US 20130141576A1 · Lord et al. · 2013 [cited by applicant]
US 20130253816A1 · Caminiti et al. · 2013 [cited by applicant]
US 20140037138A1 · Sato · 2014 [cited by examiner]
US 20140195138A1 · Stelzig et al. · 2014 [cited by applicant]
US 20140244198A1 · Mayer · 2014 [cited by applicant]
US 20150302756A1 · Guehring et al. · 2015 [cited by applicant]
US 20150316928A1 · Guehring et al. · 2015 [cited by applicant]
US 20150321337A1 · Stephens, Jr. · 2015 [cited by applicant]
US 20160005333A1 · Naouri · 2016 [cited by applicant]
US 20160209840A1 · Kim · 2016 [cited by applicant]
US 20160314224A1 · Wei et al. · 2016 [cited by applicant]
US 20170031361A1 · Olson · 2017 [cited by examiner]
US 20170072853A1 · Matsuoka et al. · 2017 [cited by applicant]
US 20170268896A1 · Bai et al. · 2017 [cited by applicant]
US 20170301111A1 · Zhao et al. · 2017 [cited by applicant]
US 20170356748A1 · Iagnemma · 2017 [cited by applicant]
US 20180011485A1 · Ferren · 2018 [cited by applicant]
US 20180053102A1 · Martinson et al. · 2018 [cited by applicant]
US 20180082596A1 · Whitlow · 2018 [cited by applicant]
US 20180089563A1 · Redding et al. · 2018 [cited by applicant]
US 20180183873A1 · Wang et al. · 2018 [cited by applicant]
US 20180184352A1 · Lopes et al. · 2018 [cited by applicant]
US 20180220283A1 · Condeixa et al. · 2018 [cited by applicant]
US 20180268281A1 · Olson et al. · 2018 [cited by applicant]
US 20180293537A1 · Kwok · 2018 [cited by applicant]
US 20180365908A1 · Liu et al. · 2018 [cited by applicant]
US 20180367997A1 · Shaw et al. · 2018 [cited by applicant]
US 20190039545A1 · Kumar et al. · 2019 [cited by applicant]
US 20190066399A1 · Jiang et al. · 2019 [cited by applicant]
US 20190106117A1 · Goldberg · 2019 [cited by applicant]
US 20190138524A1 · Singh et al. · 2019 [cited by applicant]
US 20190196465A1 · Hummelshoj · 2019 [cited by applicant]
US 20190220011A1 · Della Penna · 2019 [cited by applicant]
US 20190236950A1 · Li et al. · 2019 [cited by applicant]
US 20190258246A1 · Liu et al. · 2019 [cited by applicant]
US 20190265059A1 · Warnick et al. · 2019 [cited by applicant]
US 20200020226A1 · Stenneth et al. · 2020 [cited by applicant]
US 20200098269A1 · Wray et al. · 2020 [cited by applicant]
US 20200122830A1 · Anderson et al. · 2020 [cited by applicant]
US 20200124447A1 · Schwindt et al. · 2020 [cited by applicant]
US 20200159227A1 · Cohen et al. · 2020 [cited by applicant]
US 20200189731A1 · Mistry et al. · 2020 [cited by applicant]
US 20200233060A1 · Lull et al. · 2020 [cited by applicant]
US 20200334762A1 · Carver · 2020 [cited by examiner]
US 20200346666A1 · Wray et al. · 2020 [cited by applicant]
US 20200400781A1 · Voorheis et al. · 2020 [cited by applicant]
US 20210116907A1 · Altman · 2021 [cited by applicant]
WO WO2015160900A1 · 2015 [cited by applicant]
D. Mehta et al “Fast Discovery Of Influential Outcomes For Risk-Aware MPDM”, Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) (2017). [cited by applicant]
D. Mehta et al , “Autonomous Navigation in Dynamic Social Environments Using Multi-Policy Decision Making”, Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2016). [cited by applicant]
A. Cunningham, et al, “MPDM: Multipolicy Decision-Making in Dynamic, Uncertain Environments for Autonomous Driving” Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) (2015). [cited by applicant]
J. Straub, “Comparing The Effect Of Pruning On A Best Path and a Naive-approach Bloackboard Solver”, International Journal of Automation and Computing (Oct. 2015). [cited by applicant]
B. Paden et al., “A Survey Of MOtion Planning And Control Techniques For Self-driving Urban Vehicles”, IEEE Transactions on Intelligent Vehicles, vol. 1, Iss. 1, (Jun. 13, 2016). [cited by applicant]
International Search Report and Written Opinion of the ISA dated Oct. 15, 2019 for PCT/US19/42235. [cited by applicant]
Neumeier Stefan, et al., “Towards a Driver Support System for Teleoperated Driving”, 2019 IEEE Intelligent Transportation Systems Conference (ITSC) (Year: 2019). [cited by applicant]
Wuthishuwong, Chairit, et al., “Vehicle to Infrastructure based Safe Trajectory Planning for Autonomous Intersection Management”, 2013 13th International Conference on ITS Telecommunications (ITST) (Year: 2013). [cited by applicant]
Japanese Office Action regarding Application No. 201955066.7, mailed Jan. 21, 2022. [cited by applicant]