IP Library › Granted Patent US 12,731,037
Granted Patent B2
US 12,731,037 · App. 18/792,095 · Granted Sep 8, 2026

Vehicle operation with machine learning

Inventors: Ding Zhao (Pittsburgh, PA); Haohong Lin (Pittsburgh, PA); Yuming Niu (Northville, MI); Wenhao Ding (San Jose, CA); Zuxin Liu (Mountain View, CA); Kalpak Kalvit (Milpitas, CA)
Assignee: Ford Global Technologies, LLC
G06N3/092B60W50/0097
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,037
App. No.
18/792,095
Filed
Aug 1, 2024
Granted
Sep 8, 2026
Kind
B2
Art Unit
3669
USPC
701/117
Abstract

A computer that includes a processor and a memory, the memory including instructions executable by the processor to operate a system based on predictions output from the machine learning system including predicted states, actions, rewards, and costs, wherein the machine learning system includes a first transformer and a second transformer and is trained based on bisimulation offline reinforcement learning, wherein the first transformer and the second transformer are based on a Markov decision process that includes the states, the actions, the rewards, and the costs. The bisimulation offline reinforcement learning can include inputting a first sequences of training states, actions, rewards, and costs to the first transformer and a second sequence of the training states, actions, rewards, and costs to the second transformer to determine bisimulation learning objectives based on latent variables output from the first transformer and the second transformer.

Claims (22)

1 . A system, comprising: a computer that includes a processor and a memory, the memory including instructions executable by the processor to: operate the system based on output from a machine learning system including predicted states, actions, rewards, and costs, wherein the machine learning system includes a first transformer and a second transformer and is trained based on bisimulation offline reinforcement learning, wherein the first transformer and the second transformer are based on a Markov decision process that includes the states, the actions, the rewards, and the costs; and wherein the bisimulation offline reinforcement learning includes inputting a first sequence of training states, actions, rewards, and costs to the first transformer and a second sequence of the training states, actions, rewards, and costs to the second transformer to determine bisimulation learning objectives based on minimizing differences between latent variables including transition dynamics distributions, rewards, and costs output from the first transformer and latent variables including transition dynamics distributions, rewards, and costs output from the second transformer, respectively.

2 . The system of claim 1 , wherein the system includes a vehicle and operating the vehicle includes determining a vehicle trajectory based on the predicted states, actions, rewards, and costs output by the machine learning system.

3 . The system of claim 1 , wherein the Markov decision process is a constrained contextual Markov decision process that includes multiple sets of the states, the actions, the rewards, and the costs included in a time sequence.

4 . The system of claim 3 , wherein the Markov decision process includes the transition dynamics distributions and a discount factor.

5 . The system of claim 1 , wherein training the machine learning system based on the bisimulation offline reinforcement learning includes minimizing the bisimulation learning objectives based on the rewards, the costs, and the transition dynamics distributions included in the latent variables from the first transformer and the second transformer that includes a Lagrangian multiplier for the costs and a 2-Wasserstein distance for the transition dynamics distributions.

6 . The system of claim 1 , wherein the rewards are based on one or more of a vehicle longitudinal direction, a vehicle speed and a vehicle goal.

7 . The system of claim 1 , wherein the costs are based on one or more of not contacting objects including other vehicles, staying on a roadway, and maintaining an upper limit on vehicle speed.

8 . The system of claim 1 , wherein the states are based on a disjoint state space that includes a video image, a lidar image, and a bird's-eye view image.

9 . The system of claim 8 , wherein the bird's-eye view image is determined based on the video image and the lidar image.

10 . The system of claim 1 , wherein the first transformer and the second transformer transform the states, the actions, the rewards, and the costs to the predicted state based on encoding the states, the actions, the rewards, and the costs to a multi-dimensional vector, applying multi-head attention included in a decoder to the multi-dimensional vector to generate latent variables, and inputting the latent variables to an encoder that generates an output prediction.

11 . The system of claim 1 , wherein the sequences of the training states, the training actions, the training rewards, and the training costs used to train the first transformer and the second transformer are based on recorded real world data.

12 . A method, comprising:

operating a system based on output from a machine learning system including predicted states, actions, rewards, and costs, wherein the machine learning system includes a first transformer and a second transformer and is trained based on bisimulation offline reinforcement learning, wherein the first transformer and the second transformer are based on a Markov decision process that includes the states, the actions, the rewards, and the costs; and

wherein the bisimulation offline reinforcement learning includes inputting a first sequences of training states, actions, rewards, and costs to the first transformer and a second sequence of the training states, actions, rewards, and costs to the second transformer to determine bisimulation learning objectives based on minimizing differences between latent variables including transition dynamics distributions, rewards, and costs output from the first transformer and latent variables including transition dynamics distributions, rewards, and costs output from the second transformer, respectively.

13 . The method of claim 12 , wherein the system includes a vehicle and operating the vehicle includes determining a vehicle trajectory based on the predicted states, actions, rewards, and costs output by the machine learning system.

14 . The method of claim 12 , wherein the Markov decision process is a constrained contextual Markov decision process that includes multiple sets of the states, the actions, the rewards, and the costs included in a time sequence.

15 . The method of claim 14 , wherein the Markov decision process includes the transition dynamics distributions and a discount factor.

16 . The method of claim 12 , wherein training the machine learning system based on the bisimulation offline reinforcement learning includes minimizing the bisimulation learning objectives based on the rewards, the costs, and the transition dynamics distributions included in the latent variables from the first transformer and the second transformer that includes a Lagrangian multiplier for the costs and a 2-Wasserstein distance for the transition dynamics distributions.

17 . The method of claim 12 , wherein the rewards are based on one or more of a vehicle longitudinal direction, a vehicle speed and a vehicle goal.

18 . The method of claim 12 , wherein the costs are based on one or more of not contacting objects including other vehicles, staying on a roadway, and maintaining an upper limit on vehicle speed.

19 . The method of claim 12 , wherein the states are based on a disjoint state space that includes a video image, a lidar image, and a bird's-eye view image.

20 . The method of claim 19 , wherein the bird's-eye view image is determined based on the video image and the lidar image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2024
From: ZHAO, DING; LIN, HAOHONG; NIU, YUMING; DING, WENHAO; LIU, ZUXIN; KALVIT, KALPAK
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 068156/0897 →
Continuity (1)
Related Publication 20260037821A1 · Feb 5, 2026
References Cited (9)
US 10726279B1 · Kim et al. · 2020 [cited by applicant]
US 11010668B2 · Kim et al. · 2021 [cited by applicant]
US 20200311613A1 · Ma · 2020 [cited by examiner]
US 20220055689A1 · Mandlekar · 2022 [cited by examiner]
CN 109709956B · 2021 [cited by applicant]
Liu X. et al., Constrained Decision Transformer for Offline Safe Reinforcement Learning, Jun. 21, 2023, pp. 1-6 (Year: 2023). [cited by examiner]
Huang Z. et al., Augmenting Reinforcement Learning With Transformer-Based Scene Representation Learning for Decision-Making of Autonomous Driving, Mar. 3, 2024, IEEE Transactions on Intelligent Vehicles, vol. 9, pp. 440… [cited by examiner]
Jang H. et al., Dynamic Occupancy Grid Map with Semantic Information Using Deep Learning-Based BEVFusion Method with Camera and LiDAR Fusion, Apr. 29, 2024, MDPI, pp. 1-15 (Year: 2024). [cited by examiner]
Shi, T., et al., “Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement,” arXiv:2110.07067v2 [cs.RO] Nov. 2, 2021, 9 pages. [cited by applicant]