IP Library Granted Patent US 12,509,072
Granted Patent B2
US 12,509,072 · App. 17/871,628 · Granted Dec 30, 2025

Task-informed behavior planning

Inventors: Xin Huang (Cambridge, MA); Guy Rosman (Newton, MA); Ashkan Mohammadzadeh Jasour (Merced, CA); Stephen G. McGill, Jr. (Cambridge, MA); John J. Leonard (Newton, MA); Brian C. Williams (Cambridge, MA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA; MASSACHUSETTS INSTITUTE OF TECHNOLOGY
B60W30/0956B60W50/14B60W60/0027B60W2554/802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,509,072
App. No.
17/871,628
Granted
Dec 30, 2025
Kind
B2
Abstract

A method for task-informed planning by a behavior planning system of a vehicle includes observing a previous trajectory of an agent within a distance from the vehicle. The method also includes predicting, by the behavior planning system, a set of potential trajectories for the agent and/or the vehicle based on observing the previous trajectory. The method further includes selecting, by the behavior planning system, a potential action from a set of potential actions associated with a task to be performed by the vehicle, each potential action being associated with a utility value based on the respective potential action and the set of potential trajectories, the selected potential action being associated with a highest utility value of respective utility values associated with the set of potential actions. The method still further includes controlling the vehicle to perform an action associated with the potential action selected by the behavior planning system.

Claims (68)

1 . A method for task-informed planning by a behavior planning system of an autonomous vehicle, comprising:

observing, via one or more sensors associated with the autonomous vehicle, a previous trajectory of an agent that is within a distance from the autonomous vehicle;

predicting, by the behavior planning system, a first set of potential trajectories for the agent and a second set of potential trajectories of the autonomous vehicle based on observing the previous trajectory;

selecting, by the behavior planning system, a potential action from a set of potential actions associated with a task to be performed by the vehicle, each potential action being associated with a utility value that is a function of both an efficiency term and a safety term, each of the efficiency term and the safety term being associated with the potential action in accordance with the first set of potential trajectories and the second set of potential trajectories, the selected potential action being associated with a highest utility value of respective utility values associated with the set of potential actions; and

autonomously performing, by the autonomous vehicle, an action associated with the selected potential action.

2 . The method of claim 1 , further comprising receiving a set of inputs associated with the task.

3 . The method of claim 2 , wherein:

the task is trajectory planning for the vehicle;

the set of inputs includes the set of potential actions;

the set of potential actions include a set of candidate trajectories of the vehicle; and

the predicted set of potential trajectories includes potential trajectories of the agent.

4 . The method of claim 3 , wherein:

the behavior planning system is trained to determine the utility value based on the function that uses the efficiency term and the safety term;

the efficiency term is based on a distance traveled by one candidate trajectory of the set of candidate trajectories; and

the safety term is based on an expected closest distance between one candidate trajectory of the set of candidate trajectories and the set of potential trajectories.

5 . The method of claim 1 , wherein:

the task is warning generation at the vehicle;

the set of potential actions include a first potential action associated with generating a warning and a second potential action associated with not generating the warning; and

the predicted set of potential trajectories include a set of potential agent trajectories a set of potential vehicle trajectories.

6 . The method of claim 5 , wherein the warning term is associated with a likelihood of a collision between each potential agent trajectory of the set of potential agent trajectories and each potential vehicle trajectory of the set of potential vehicle trajectories.

7 . The method of claim 1 , further comprising:

training the behavior planning system to predict the set of potential trajectories by minimizing a loss between a set of potential training trajectories and a ground truth trajectory; and

training the behavior planning system to select the potential action by minimizing a cross entropy between a decision utility and a ground truth decision.

8 . An apparatus for task-informed planning by a behavior planning system of a vehicle, comprising:

at least one processor; and

at least one memory coupled with the at least one processor and storing instructions operable, when executed by the at least one processor, to cause the apparatus:

observe, via one or more sensors associated with the autonomous vehicle, a previous trajectory of an agent that is within a distance from the autonomous vehicle;

predict, by the behavior planning system, a first set of potential trajectories for the agent and a second set of potential trajectories of the autonomous vehicle based on observing the previous trajectory;

select, by the behavior planning system, a potential action from a set of potential actions associated with a task to be performed by the vehicle, each potential action being associated with a utility value that is a function of both an efficiency term and a safety term, each of the efficiency term and the safety term being associated with the potential action in accordance with the first set of potential trajectories and the second set of potential trajectories, the selected potential action being associated with a highest utility value of respective utility values associated with the set of potential actions; and

autonomously perform, by the autonomous vehicle, an action associated with the selected potential action.

9 . The apparatus of claim 8 , wherein execution of the instructions further cause the apparatus to receive a set of inputs associated with the task.

10 . The apparatus of claim 9 , wherein:

the task is trajectory planning for the vehicle;

the set of inputs includes the set of potential actions;

the set of potential actions include a set of candidate trajectories of the vehicle; and

the predicted set of potential trajectories includes potential trajectories of the agent.

11 . The apparatus of claim 10 , wherein:

the behavior planning system is trained to determine the utility value based on the function that uses the efficiency term and the safety term;

the efficiency term is based on a distance traveled by one candidate trajectory of the set of candidate trajectories; and

the safety term is based on an expected closest distance between one candidate trajectory of the set of candidate trajectories and the set of potential trajectories.

12 . The apparatus of claim 8 , wherein:

the task is warning generation at the vehicle;

the set of potential actions include a first potential action associated with generating a warning and a second potential action associated with not generating the warning; and

the predicted set of potential trajectories include a set of potential agent trajectories a set of potential vehicle trajectories.

13 . The apparatus of claim 12 , wherein the warning term is associated with a likelihood of a collision between each potential agent trajectory of the set of potential agent trajectories and each potential vehicle trajectory of the set of potential vehicle trajectories.

14 . The apparatus of claim 8 , wherein execution of the instructions further cause the apparatus to:

train the behavior planning system to predict the set of potential trajectories by minimizing a loss between a set of potential training trajectories and a ground truth trajectory; and

train the behavior planning system to select the potential action by minimizing a cross entropy between a decision utility and a ground truth decision.

15 . A non-transitory computer-readable medium having program code recorded thereon for task-informed planning by a behavior planning system of a vehicle, the program code executed by at least one processor and comprising:

program code to observe, via one or more sensors associated with the autonomous vehicle, a previous trajectory of an agent that is within a distance from the autonomous vehicle;

program code to predict, by the behavior planning system, a first set of potential trajectories for the agent and a second set of potential trajectories of the autonomous vehicle based on observing the previous trajectory;

program code to select, by the behavior planning system, a potential action from a set of potential actions associated with a task to be performed by the vehicle, each potential action being associated with a utility value that is a function of both an efficiency term and a safety term, each of the efficiency term and the safety term being associated with the potential action in accordance with the first set of potential trajectories and the second set of potential trajectories, the selected potential action being associated with a highest utility value of respective utility values associated with the set of potential actions; and

program code to autonomously perform, by the autonomous vehicle, an action associated with the selected potential action.

16 . The non-transitory computer-readable medium of claim 15 , wherein the program code further comprises program code to receive a set of inputs associated with the task.

17 . The non-transitory computer-readable medium of claim 16 , wherein:

the task is trajectory planning for the vehicle;

the set of inputs includes the set of potential actions;

the set of potential actions include a set of candidate trajectories of the vehicle; and

the predicted set of potential trajectories includes potential trajectories of the agent.

18 . The non-transitory computer-readable medium of claim 17 , wherein:

the behavior planning system is trained to determine the utility value based on the function that uses the efficiency term and the safety term;

the efficiency term is based on a distance traveled by one candidate trajectory of the set of candidate trajectories; and

the safety term is based on an expected closest distance between one candidate trajectory of the set of candidate trajectories and the set of potential trajectories.

19 . The non-transitory computer-readable medium of claim 15 , wherein:

the task is warning generation at the vehicle;

the set of potential actions include a first potential action associated with generating a warning and a second potential action associated with not generating the warning; and

the predicted set of potential trajectories include a set of potential agent trajectories a set of potential vehicle trajectories.

20 . The non-transitory computer-readable medium of claim 19 , wherein the warning term is associated with a likelihood of a collision between each potential agent trajectory of the set of potential agent trajectories and each potential vehicle trajectory of the set of potential vehicle trajectories.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2023
From: ROSMAN, GUY; MCGILL, STEPHEN G., JR.; LEONARD, JOHN J.
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 065428/0396 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2023
From: HUANG, XIN; JASOUR, ASHKAN MOHAMMADZADEH; WILLIAMS, BRIAN C.
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 065428/0729 →
Continuity (2)
Provisional Application 63243492 · Sep 13, 2021
Related Publication 20230085422A1 · Mar 16, 2023
References Cited (12)
US 11345342B2 · Gutierrez · 2022 [cited by examiner]
US 11851081B2 · Refaat · 2023 [cited by examiner]
US 11981349B2 · Nister · 2024 [cited by examiner]
US 20180141544A1 · Xiao · 2018 [cited by examiner]
US 20190382007A1 · Casas et al. · 2019 [cited by applicant]
US 20210009121A1 · Oboril · 2021 [cited by examiner]
US 20210173402A1 · Chang · 2021 [cited by examiner]
US 20210200212A1 · Urtasun · 2021 [cited by examiner]
US 20210279511A1 · Gordon et al. · 2021 [cited by applicant]
US 20210354718A1 · Lu et al. · 2021 [cited by applicant]
Huang, et al., “CARPAL: Confidence-Aware Intent Recognition for Parallel Autonomy,” IEEE Robotics and Automation Letters 2021, pp. 1-1. 10. Published online at arXiv:2003.08003 on Mar. 17, 2021. [cited by applicant]
Sadat, et al., “Jointly Learnable Behavior and Trajectory Planning for Self-Driving Vehicles,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Published online at arXiv:1910.04586 on Oct… [cited by applicant]