IP Library Granted Patent US 12,415,538
Granted Patent B2
US 12,415,538 · App. 17/576,553 · Granted Sep 16, 2025

Systems and methods for pareto domination-based learning

Inventors: Brian D. Ziebart (Chicago, IL); Paul Vernaza (Cupertino, CA)
Assignee: AURORA OPERATIONS, INC.
B60W60/001B60W50/00G05B13/0265B60W2050/0022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,415,538
App. No.
17/576,553
Granted
Sep 16, 2025
Kind
B2
Abstract

Techniques for improving the performance of an autonomous vehicle (AV) are described herein. A system can determine a plan for the AV in a driving scenario that optimizes an initial cost function of a control algorithm of the AV. The system can obtain data describing an observed human driving path in the driving scenario. Additionally, the system can determine for each cost dimension in the plurality of cost dimensions, a quantity that compares the estimated cost to the observed cost of the observed human driving path. Moreover, the system can determine a function of a sum of the quantities determined for each cost dimension in the plurality of cost dimensions. Subsequently, the system can use an optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities.

Claims (45)

1. A method for improving performance of an autonomous vehicle (AV), the method comprising:

(a) determining a plan for the AV in a driving scenario that optimizes an initial cost function of a control algorithm of the AV, wherein the initial cost function comprises a plurality of cost dimensions and a plurality of weights applied to the plurality of cost dimensions, and wherein the plan comprises a plurality of estimated costs associated with the plurality of cost dimensions;

(b) obtaining data describing an observed human driving path in the driving scenario, wherein the data comprises a first plurality of observed costs associated with the plurality of cost dimensions of the initial cost function;

(c) determining, for each cost dimension in the plurality of cost dimensions, a quantity by which the estimated cost exceeds the observed cost of the observed human driving path;

(d) determining a function of a sum of the quantities determined for each cost dimension in the plurality of cost dimensions, wherein the function of the sum of the quantities comprises a margin by which the estimated cost for each cost dimension in the plurality of cost dimensions exceeds the observed cost of the observed human driving path;

(e) using an optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities; and

(f) controlling a motion of the AV in accordance with the control algorithm of the AV. the control algorithm comprising adjustments made to the one or more weights applied to the plurality of cost dimensions of the initial cost function.

2. The method of claim 1 , wherein the margin is indicative of an expected dominance gap between the estimated cost and the observed cost.

3. The method of claim 1 , wherein the function comprises a plurality of learned parameters associated with the plurality of cost dimensions, and wherein the method further comprises, prior to (e), updating, using the optimization algorithm, the plurality of learned parameters to optimize an output of the function.

4. The method of claim 3 , wherein (e) comprises adjusting the one or more weights based at least in part on the updated plurality of learned parameters.

5. The method of claim 1 , wherein the function of the sum of the quantities comprises a respective margin slope for each cost dimension in the plurality of cost dimensions, and wherein the method further comprises:

setting a value of the respective margin slope for each cost dimension in the plurality of cost dimensions based on the plan for the AV, and wherein (e) comprises adjusting the one or more weights in the plurality of weights based on the respective margin slopes.

6. The method of claim 1 , wherein the one or more weights of the plurality of weights is adjusted to minimize the function of the sum of quantities.

7. The method of claim 1 , wherein the function of the sum of the quantities is optimized when the function of the sum of the quantities achieves a global minimum for all of the plurality of weights applied to the plurality of cost dimensions.

8. The method of claim 1 , wherein the function of the sum of the quantities is a total sum of the quantities determined for each cost dimension in the plurality of cost dimensions.

9. The method of claim 1 , wherein the plurality of cost dimensions includes a control cost, nudge lateral cost, and a lateral jerk cost.

10. The method of claim 1 , wherein the optimization algorithm comprises a machine-learned model that is trained based on the data describing the observed human driving path in the driving scenario.

11. The method of claim 1 , wherein the data further comprises a second plurality of observed costs associated with the plurality of cost dimensions of the initial cost function; and

the method further comprises:

determining the observed cost of the observed human driving path by averaging, for each cost dimension in the plurality of cost dimensions, the first plurality of observed costs with the second plurality of observed costs.

12. The method of claim 1 , wherein the determined plan for the AV includes a human-behavior prediction portion and an AV-behavior prediction portion.

13. The method of claim 12 , the method further comprising:

generating a first sparse plan distribution based on the human-behavior prediction portion of the determined plan; and

generating a second sparce plan distribution based on the AV-behavior prediction portion.

14. An autonomous vehicle control system for an autonomous vehicle (AV), the autonomous vehicle control system comprising:

one or more processors;

one or more non-transitory computer-readable media that store an optimization algorithm, wherein the optimization algorithm is configured to optimize an initial cost function of a control algorithm of an AV, wherein the initial cost function comprises a plurality of cost dimensions and a plurality of weights applied to the plurality of cost dimensions; and

instructions that, when executed by the one or more processors, cause the autonomous vehicle control system to perform operations, the operations comprising:

determining a plan for the AV in a driving scenario that optimizes the initial cost function of the control algorithm of the AV, wherein the plan comprises a plurality of estimated costs associated with the plurality of cost dimensions;

obtaining data describing an observed human driving path in the driving scenario, wherein the data comprises a plurality of observed costs associated with the plurality of cost dimensions of the initial cost function;

determining for each cost dimension in the plurality of cost dimensions, a quantity by which the estimated cost exceeds the observed cost of the observed human driving path;

determining, a function of a sum of the quantities determined for each cost dimension in plurality of cost dimensions, wherein the function of the sum of the quantities comprises a margin by which the estimated cost for each cost dimension in the plurality of cost dimensions exceeds the observed cost of the observed human driving path;

using the optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities; and

controlling a motion of the AV in accordance with the control algorithm of the AV, the control algorithm comprising adjustments made to the one or more weights applied to the plurality of cost dimensions of the initial cost function.

15. The autonomous vehicle control system of claim 14 , wherein the estimated cost for each cost dimension in the plurality of cost dimensions exceeds the observed cost of the observed human driving path by at least a margin, the operations further comprising:

updating, using the optimization algorithm, the margin based on the adjusted one or more weights of the plurality of weights applied to the plurality of cost dimensions.

16. One or more non-transitory computer-readable media that store:

an optimization algorithm, wherein the optimization algorithm is configured to optimize an initial cost function of a control algorithm of an autonomous vehicle (AV), wherein the initial cost function comprises a plurality of cost dimensions and a plurality of weights applied to the plurality of cost dimensions; and

instructions that, when executed by one or more processors, cause an autonomous vehicle control system to perform operations, the operations comprising:

determining a plan for the AV in a driving scenario that optimizes the initial cost function of the control algorithm of the AV, wherein the plan comprises a plurality of estimated costs associated with the plurality of cost dimensions,

obtaining data describing an observed human driving path in the driving scenario, wherein the data comprises a plurality of observed costs associated with the plurality of cost dimensions of the initial cost function;

determining for each cost dimension in the plurality of cost dimensions, a quantity by which the estimated cost exceeds the observed cost of the observed human driving path;

determining, a function of a sum of the quantities determined for each cost dimension in plurality of cost dimensions, wherein the function of the sum of the quantities comprises a margin by which the estimated cost for each cost dimension in the plurality of cost dimensions exceeds the observed cost of the observed human driving path;

using the optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities; and

controlling a motion of the AV in accordance with the control algorithm of the AV, the control algorithm comprising adjustments made to the one or more weights applied to the plurality of cost dimensions of the initial cost function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2022
From: ZIEBART, BRIAN D.; VERNAZA, PAUL
To: AURORA OPERATIONS, INC.
Reel/Frame 059455/0415 →
Continuity (1)
Related Publication 20230227061A1 · Jul 20, 2023
References Cited (31)
US 20060178824A1 · Ibrahim · 2006 [cited by applicant]
US 20100106603A1 · Dey · 2010 [cited by examiner]
US 20170168485A1 · Berntorp · 2017 [cited by examiner]
US 20190180115A1 · Zou · 2019 [cited by examiner]
US 20190196419A1 · Wee · 2019 [cited by examiner]
US 20200139959A1 · Akella · 2020 [cited by examiner]
US 20200159225A1 · Zeng · 2020 [cited by examiner]
US 20200216094A1 · Zhu · 2020 [cited by examiner]
US 20210020045A1 · Huang · 2021 [cited by examiner]
US 20210200212A1 · Urtasun et al. · 2021 [cited by applicant]
US 20210264287A1 · Gardner · 2021 [cited by examiner]
US 20210291862A1 · Jiang · 2021 [cited by examiner]
US 20210403034A1 · Lapin et al. · 2021 [cited by applicant]
US 20210403045A1 · Lin et al. · 2021 [cited by applicant]
US 20220048187A1 · Rudenko · 2022 [cited by examiner]
DE 102019216232A1 · 2021 [cited by examiner]
KR 102223347 · 2021 [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2023/010167, mailed May 1, 2023, 7 pages. [cited by applicant]
Abbeel et al, “Apprenticeship Learning via Inverse Reinforcement Learning”, Conference on Machine Learning, Jul. 4-8, 2004, Banff Alberta, Canada, 8 pages. [cited by applicant]
Brown et al, “Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations”, International Conference on Machine Learning, Jun. 10-15, 2019, Long Beach, California, United States, … [cited by applicant]
Crammer et al, “On the Algorithmic Implementation of Multiclass Kernel-Based Vector Machines”, Journal of Machine Learning Research, vol. 2, Dec. 2001, pp. 265-292. [cited by applicant]
Doerr et al, “Direct Loss Minimization Inverse Optimal Control”, In Robotics: Science and Systems, 2015, 9 pages. [cited by applicant]
Liu, “Fisher Consistency of Multicategory Support Vector Machines”, Artificial Intelligence and Statistics, 8 pages. [cited by applicant]
Ng et al, “Algorithms for Inverse Reinforcement Learning” International Conference on Machine Learning, Jun. 29-Jul. 2, 2000, Stanford, California, United States, 8 pages, 2000. [cited by applicant]
Ratliff et al, “Maximum Margin Planning”, International Conference on Machine Learning, Jun. 25-29, 2006, Pittsburgh, Pennsylvania, United States, 8 pages. [cited by applicant]
Syed et al, “A Game-Theoretic Approach to Apprenticeship Learning”, Conference on Neural Information Processing Systems, Dec. 8-10, 2008, British Columbia, Canada, 8 pages. [cited by applicant]
Taskar et al, “Max-Margin Markov Networks”, Conference on Neural Information Processing Systems, Dec. 8-13, 2003, Vancouver, Canada. 8 pages. [cited by applicant]
Tsochantaridis et al, “Support Vector Machine Learning for Interdependent and Structured Output Spaces”, International Conference on Machine Learning, Aug. 21-24, 2003, Washington, District of Columbia, United States, 8… [cited by applicant]
Vapnik et al, “Bounds on Error Expectation for Support Vector Machines”, Neural Computation, vol. 12 No. 9, 2000, pp. 2013-2036. [cited by applicant]
Ziebart et al, “Maximum Entropy Inverse Reinforcement Learning”, National Conference on Artificial Intelligence, Jul. 13-17, 2008, Chicago, Illinois, United States, pp. 1433-1438. [cited by applicant]
Ziebart et al, “Probabilistic Pointing Target Prediction via Inverse Optimal Control”, International Conference on Intelligent User Interfaces, Feb. 14-17, 2012, Lisbon, Portugal, 10 pages. [cited by applicant]