IP Library › Granted Patent US 12,528,507
Granted Patent B2
US 12,528,507 · App. 18/388,606 · Granted Jan 20, 2026

Systems and methods for vision-language planning (VLP) foundation models for autonomous driving

Inventors: Chenbin Pan (Syracuse, NY); Burhaneddin Yaman (San Jose, CA); Tommaso Nesti (Mountain View, CA); Abhirup Mallik (Santa Clara, CA); Yuliang Guo (Palo Alto, CA); Liu Ren (Saratoga, CA)
Assignee: Robert Bosch GmbH
B60W60/0015G06V20/56B60W2420/403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,528,507
App. No.
18/388,606
Granted
Jan 20, 2026
Kind
B2
Abstract

Methods and systems for training an autonomous driving system using a vision-language planning (VLP) model. Image data is obtained from a vehicle-mounted camera, encompassing details about agents situated within the external environment. Via image processing, the system identifies these agents within the environment. A Bird's Eye View (BEV) representation of the surroundings is then generated, encapsulating the spatiotemporal information linked to the vehicle and the recognized agents. Execution of the VLP machine learning model begins by extracting vision-based planning features from the BEV, and receiving or generating textual information characterizing various attributes of the vehicle within the environment. Text-based planning features are extracted from this textual information. To enhance model performance, a contrastive learning model is engaged to establish similarities between the vision-based and text-based planning features, and a predicted trajectory is output based on the similarities.

Claims (58)

1 . A method of training an autonomous driving system utilizing a vision-language planning (VLP) machine learning model, the method comprising:

receiving image data generated from a camera mounted to a vehicle, wherein the image data includes agents in an environment outside the vehicle;

via image processing, detecting the agents in the environment based on the image data;

generating a bird eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with the vehicle and the detected agents; and

executing a vision-language planning (VLP) machine learning model to:

extract vision-based planning features from the BEV, wherein the vision-based planning features include the spatiotemporal information associated with the vehicle,

generate text information associated with the environment, wherein the text information describes qualities of the vehicle in the environment;

extract text-based planning features from the text information,

execute a contrastive learning model to derive similarities between the vision-based planning features and the text-based planning features, and

generate a predicted trajectory of the vehicle based on the similarities.

2 . The method of claim 1 , further comprising:

determining a loss between the predicted trajectory of the vehicle and a ground truth trajectory of the vehicle; and

repeat the steps of claim 1 until convergence to minimize the loss.

3 . The method of claim 1 , wherein the text information is generated using a template and ground truth information associated with the environment existent in training data.

4 . The method of claim 1 , wherein an output of the contrastive learning is used as a training loss for training the VLP machine learning model.

5 . The method of claim 1 , wherein the contrastive learning model includes:

a text encoder configured to output a text-based vector representing text-based features associated with the text information of the environment; and

an image encoder configured to output an image-based vector representing image-based features associated with the detected agents in the BEV.

6 . The method of claim 5 , wherein the contrastive learning model is further configured to execute a dot product to evaluate similarities between the text-based vector and the image-based vector.

7 . The method of claim 1 , wherein the contrastive learning model is further configured to push apart dissimilarities between the vision-based planning features and the text-based planning features.

8 . The method of claim 1 , wherein the agents in the environment include at least one of a pedestrian, another vehicle, or a cyclist.

9 . A system utilizing a vision-language planning (VLP) machine learning model, the system comprising:

a camera mounted to a vehicle and configured to generate image data associated with agents in an environment outside the vehicle;

a processor; and

memory including instructions that, when executed by the processor, cause the processor to:

process the image data to detect the agents in the environment,

generate a bird eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with the vehicle and the detected agents, and

execute a vision-language planning (VLP) machine learning model to:

extract vision-based planning features from the BEV, wherein the vision-based planning features include the at least some of the spatiotemporal information associated with the vehicle,

receive text information associated with the environment, wherein the text information describes qualities of the vehicle in the environment,

extract text-based planning features from the text information,

execute a contrastive learning model to derive similarities between the vision-based planning features and the text-based planning features, and

generate a predicted trajectory of the vehicle based on the similarities.

10 . The system of claim 9 , wherein the memory includes further instructions that, when executed by the processor, cause the processor to:

determine a loss between the predicted trajectory of the vehicle and a ground truth trajectory of the vehicle; and

execute the VLP model until convergence to minimize the loss.

11 . The system of claim 9 , wherein the text information is generated using a template and ground truth information associated with the environment existent in training data.

12 . The system of claim 9 , wherein an output of the contrastive learning is used as a training loss for training the VLP machine learning model.

13 . The system of claim 9 , wherein the contrastive learning model includes:

a text encoder configured to output a text-based vector representing text-based features associated with the text information of the environment; and

an image encoder configured to output an image-based vector representing image-based features associated with the detected agents in the BEV.

14 . The system of claim 13 , wherein the contrastive learning model is further configured to execute a dot product to evaluate similarities between the text-based vector and the image-based vector.

15 . The system of claim 9 , wherein the contrastive learning model is further configured to push apart dissimilarities between the vision-based planning features and the text-based planning features.

16 . The system of claim 9 , wherein the agents in the environment include at least one of a pedestrian, another vehicle, or a cyclist.

17 . A method of training an autonomous driving system, the method comprising:

receiving image data generated from a camera mounted to a vehicle, wherein the image data includes agents in an environment outside the vehicle;

generating a bird eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with the vehicle and the agents;

based on the BEV, executing a perception model to detect the agents in the environment and associated information about the detected agents;

based on the BEV, executing a prediction model to estimate trajectories of the detected agents;

based on the BEV, executing a vision-language planning (VLP) model to output a predicted trajectory of the vehicle, wherein the VLP model is configured to:

extract vision-based planning features from the BEV, wherein the vision-based planning features include the spatiotemporal information associated with the vehicle,

receive text information associated with the environment, wherein the text information describes qualities of one or more of the agents in the environment,

extract text-based planning features from the text information,

perform contrastive learning to derive similarities between the vision-based planning features and the text-based planning features, and

output the predicted trajectory based on the similarities.

18 . The method of claim 17 , wherein the text information is generated using a template and ground truth information associated with the environment existent in training data.

19 . The method of claim 17 , wherein an output of the contrastive learning is used as a training loss for training the VLP machine learning model.

20 . The method of claim 17 , wherein the agents in the environment include at least one of a pedestrian, another vehicle, or a cyclist.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: PAN, CHENBIN; YAMAN, BURHANEDDIN; NESTI, TOMMASO; MALLIK, ABHIRUP; GUO, YULIANG; REN, LIU
To: ROBERT BOSCH GMBH
Reel/Frame 067065/0659 →
Continuity (1)
Related Publication 20250153736A1 · May 15, 2025
References Cited (22)
US 12158762B1 · O'Hara · 2024 [cited by examiner]
US 20220230016A1 · Ahamed et al. · 2022 [cited by applicant]
US 20220332317A1 · Lewandowski et al. · 2022 [cited by applicant]
US 20230252795A1 · Tong · 2023 [cited by examiner]
US 20230278587A1 · Wu · 2023 [cited by applicant]
US 20240087136A1 · Waez et al. · 2024 [cited by applicant]
US 20240300490A1 · Yang et al. · 2024 [cited by applicant]
CN 111002980A · 2020 [cited by applicant]
CN 114071013B · 2023 [cited by applicant]
DE 102020202476A1 · 2021 [cited by applicant]
Learning Transferable Visual Models From Natural Language Supervision, to Radford et al. pp. 1-48 published Feb. 2021 (Year: 2021). [cited by examiner]
Xiao et al., “APPLD: Adaptive Planner Parameter Learning From Demonstration,” IEEE Robotics and Automation Letters, vol. 5, No. 3, Jul. 2020, 7 pages. [cited by applicant]
Larson et al., “Derivative-free optimization methods,” arXiv:1904.11585v2 [math.OC] Jun. 25, 2019, 94 pages. [cited by applicant]
Hu et al., “Planning-oriented Autonomous Driving,” Computer Vision Foundation, 10 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021, 16 pages. [cited by applicant]
Jaime Fernandez, “Error Metrics for Trajectory Prediction Accuracy”, Feb. 7, 2019, Wordpress (Year: 2019). [cited by applicant]
Mrinal Kalakrishnan et al., “STOMP: Stochastic Trajectory Optimization for Motion Planning”, May 13, 2011, IEEE, pp. 4569-4574 (Year: 2011). [cited by applicant]
Yukang Kou and Changxi Ma, “Dual-objective intelligent vehicle lane changing trajectory planning based on polynomial optimization”, Mar. 9, 2023, Elsevier, No. 128665, pp. 1-16 (Year: 2023). [cited by applicant]
Sampurna Mandal et al., “Motion Prediction for Autonomous Vehicles from Lyft Dataset using Deep Learning”, Oct. 31, 2010, IEEE, 5th ICCCA, pp. 768-773 (Year: 2010). [cited by applicant]
Nathan Ratliff et al., “CHOMP: Gradient Optimization Techniques for Efficient Motion Planning”, May 17, 2009, IEEE, pp. 489-494 (Year: 2009). [cited by applicant]
Rahel Vortmeyer-Kley et al., “A trajectory based loss function to learn missing terms in bifurcating dynamical systems”, 2011, Scientific Reports, pp. 1-13 (Year: 2011). [cited by applicant]
Olov Andersson et al., “Model-Predictive Control with Stochastic Collision Avoidance using Bayesian Policy Optimization”, May 21, 2016, IEEE, 2016 ICRA, pp. 4597-4604 (Year: 2016), 8 Pages. [cited by applicant]