IP Library Granted Patent US 12,515,648
Granted Patent B2
US 12,515,648 · App. 18/110,082 · Granted Jan 6, 2026

System and method for training a policy using closed-loop weighted empirical risk minimization

Inventors: Eesha Kumar (London, GB); Yiming Zhang (Palo Alto, CA); Stefano Pini (London, GB); Simon A. I. Stent (Cambridge, MA); Ana Sofia Rufino Ferreira (Berkeley, CA); Sergey Zagoruyko (London, GB); Christian Samuel Perone (London, GB)
Assignee: Woven By Toyota, Inc.
B60W30/095B60W60/0027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,515,648
App. No.
18/110,082
Granted
Jan 6, 2026
Kind
B2
Abstract

Systems and methods for training a policy are disclosed. In one example, a system includes a processor and a memory with instructions that cause the processor to train the policy using a training data set with training scenes to generate an identification policy and perform a closed-loop simulation on the identification policy to collect closed-loop metrics. Based on the closed-loop metrics, the instructions cause the processor to construct an error set of the training scenes and construct an upsampled training set by upsampling the error set. After that, the policy is trained using the upsampled training set to generate a final policy.

Claims (40)

1 . A method for training a policy comprising steps of:

training the policy using a training data set having training scenes to generate an identification policy;

performing a closed-loop simulation on the identification policy to collect closed-loop metrics that count failure scenes where an agent executing the policy violated a constraint;

based on the closed-loop metrics, constructing an error set that includes the failure scenes;

constructing an upsampled training set by upsampling the error set;

training the policy using the upsampled training set to generate a final policy; and

controlling a movement of a vehicle using the final policy.

2 . The method of claim 1 , wherein the upsampled training set includes the training data set and the error set that has been upsampled.

3 . The method of claim 1 , further comprising the step of determining which scenes during the closed-loop simulation that violated a constraint.

4 . The method of claim 3 , wherein the constraint includes at least one of collisions and distances from a reference trajectory.

5 . The method of claim 3 , wherein the error set includes the training scenes that violated the constraint.

6 . The method of claim 1 , wherein the final policy is a self-driving vehicle policy.

7 . The method of claim 1 , wherein the step of performing the closed-loop simulation includes performing rollouts of the identification policy in log-replayed scenes on a simulator.

8 . A system for training a policy, the system comprising:

a processor; and

a memory in communication with the processor with instructions that, when executed by the processor, cause the processor to:

train the policy using a training data set having training scenes to generate an identification policy,

perform a closed-loop simulation on the identification policy to collect closed-loop metrics that count failure scenes where an agent executing the policy violated a constraint,

based on the closed-loop metrics, construct an error set that includes the failure scenes,

construct an upsampled training set by upsampling the error set,

train the policy using the upsampled training set to generate a final policy; and

controlling a movement of a vehicle using the final policy.

9 . The system of claim 8 , wherein the upsampled training set includes the training data set and the error set that has been upsampled.

10 . The system of claim 8 , wherein the memory further comprises instructions that, when executed by the processor, cause the processor to determine which scenes during the closed-loop simulation that violated a constraint.

11 . The system of claim 10 , wherein the constraint includes at least one of collisions and distances from a reference trajectory.

12 . The system of claim 10 , wherein the error set includes the training scenes that violated the constraint.

13 . The system of claim 8 , wherein the final policy is a self-driving vehicle policy.

14 . The system of claim 8 , wherein the memory further comprises instructions that, when executed by the processor, cause the processor to perform rollouts of the identification policy in log-replayed scenes on a simulator.

15 . A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to:

train a policy using a training data set having training scenes to generate an identification policy;

perform a closed-loop simulation on the identification policy to collect closed-loop metrics that count failure scenes where an agent executing the policy violated a constraint;

based on the closed-loop metrics, construct an error set that includes the failure scenes;

construct an upsampled training set by upsampling the error set;

train the policy using the upsampled training set to generate a final policy; and

control a movement of a vehicle using the final policy.

16 . The non-transitory computer-readable medium of claim 15 , wherein the upsampled training set includes the training data set and the error set that has been upsampled.

17 . The non-transitory computer-readable medium of claim 15 , wherein the non-transitory computer-readable medium further comprises instructions that, when executed by the processor, cause the processor to determine which scenes during the closed-loop simulation that violated a constraint.

18 . The non-transitory computer-readable medium of claim 17 , wherein the constraint includes at least one of collisions and distances from a reference trajectory.

19 . The non-transitory computer-readable medium of claim 17 , wherein the error set includes the training scenes that violated the constraint.

20 . The non-transitory computer-readable medium of claim 15 , wherein the non-transitory computer-readable medium further comprises instructions that, when executed by the processor, cause the processor to perform rollouts of the identification policy in log-replayed scenes on a simulator.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Jun 23, 2023
From: WOVEN ALPHA, INC.; WOVEN BY TOYOTA, INC.
To: WOVEN BY TOYOTA, INC.
Reel/Frame 064044/0373 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: KUMAR, EESHA; ZHANG, YIMING; PINI, STEFANO; STENT, SIMON A.I.; FERREIRA, ANA SOFIA RUFINO; ZAGORUYKO, SERGEY; PERONE, CHRISTIAN SAMUEL
To: WOVEN ALPHA, INC.
Reel/Frame 062901/0651 →
Continuity (2)
Provisional Application 63406489 · Sep 14, 2022
Related Publication 20240092356A1 · Mar 21, 2024
References Cited (48)
US 11126180B1 · Kobilarov · 2021 [cited by examiner]
US 11741274B1 · Crego · 2023 [cited by examiner]
US 11938966B1 · Besson · 2024 [cited by examiner]
US 12071157B1 · Prioletti · 2024 [cited by examiner]
US 12103521B2 · Das · 2024 [cited by examiner]
US 12162500B1 · Egbert · 2024 [cited by examiner]
US 12175764B1 · Song · 2024 [cited by examiner]
US 20160107682A1 · Tan · 2016 [cited by examiner]
US 20170043768A1 · Prokhorov · 2017 [cited by examiner]
US 20170063599A1 · Wu · 2017 [cited by examiner]
US 20180032891A1 · Ba · 2018 [cited by examiner]
US 20200074266A1 · Peake · 2020 [cited by examiner]
US 20200110416A1 · Hong · 2020 [cited by examiner]
US 20210080969A1 · Hayes · 2021 [cited by examiner]
US 20210192748A1 · Morales Morales · 2021 [cited by examiner]
US 20220055618A1 · Toyoda · 2022 [cited by examiner]
US 20220180890A1 · Ramaiah · 2022 [cited by examiner]
US 20220204034A1 · Stein · 2022 [cited by examiner]
US 20220227367A1 · Kario · 2022 [cited by examiner]
US 20220234578A1 · Das · 2022 [cited by examiner]
US 20220247618A1 · Côté · 2022 [cited by examiner]
US 20230030104A1 · Puchkarev · 2023 [cited by examiner]
US 20230120917A1 · Muthusami · 2023 [cited by examiner]
US 20230195122A1 · Shenfeld · 2023 [cited by examiner]
US 20230326159A1 · Shayani · 2023 [cited by examiner]
US 20230326335A1 · Ding · 2023 [cited by examiner]
US 20230339526A1 · Kernwein · 2023 [cited by examiner]
US 20230351772A1 · Poltoraski · 2023 [cited by examiner]
US 20230373529A1 · Tomov · 2023 [cited by examiner]
US 20240078470A1 · Lindenau · 2024 [cited by examiner]
US 20240229732A1 · Abrosimov · 2024 [cited by examiner]
US 20240394541A1 · Cemgil · 2024 [cited by examiner]
US 20240400045A1 · Gochev · 2024 [cited by examiner]
CA 3099659A1 · 2019 [cited by examiner]
CN 114630042A · 2022 [cited by examiner]
DE 102022209278A1 · 2024 [cited by examiner]
GB 2605991A · 2022 [cited by examiner]
KR 102626145B1 · 2024 [cited by examiner]
WO WO2020224910A1 · 2020 [cited by examiner]
WO WO2020250019A1 · 2020 [cited by examiner]
WO WO2021206681A1 · 2021 [cited by examiner]
Liu, E. Z., Haghgoo, B., Chen, A. S., Raghunathan, A., Koh, P. W., Sagawa, S., Liang, P., & Finn, C. (2021). Just Train Twice: Improving Group Robustness without Training Group Information. [cited by applicant]
M. Bansal, A. Krizhevsky, A. Ogale. “ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst.” Dec. 7, 2018. arXiv:1812.03079v1. [cited by applicant]
Scheel, Oliver. Luca Bergamini, Maciej Wolczyk, Blazej Osinski, and Peter Ondruska. “Urban driver: Learning to drive from real-world demonstrations using policy gradients.” In Conference on Robot Learning, pp. 718-728. … [cited by applicant]
Idrissi, Badr Youbi, Martin Arjovsky, Mohammad Pezeshki, and David Lopez-Paz. “Simple data balancing achieves competitive worst-group-accuracy.” In Conference on Causal Learning and Reasoning, pp. 336-351. PMLR, 2022. h… [cited by applicant]
Kirichenko, Polina, Pavel Izmailov, and Andrew Gordon Wilson. “Last layer re-training is sufficient for robustness to spurious correlations.” arXiv preprint arXiv:2204.02937 (2022). https://arxiv.org/pdf/2204.02937. [cited by applicant]
Shimodaira, Hidetoshi. “Improving predictive inference under covariate shift by weighting the log-likelihood function.” Journal of statistical planning and inference 90, No. 2 (2000): 227-244. https://www.academia.edu/d… [cited by applicant]
Ross, Stephane, and Drew Bagnell. “Efficient reductions for imitation learning.” In Proceedings of the thirteenth International conference on artificial intelligence and statistics, pp. 661-668. JMLR Workshop and Confer… [cited by applicant]