IP Library Granted Patent US 11,640,562
Granted Patent B1
US 11,640,562 · App. 17/967,725 · Granted May 2, 2023

Counterexample-guided update of a motion planner

Inventors: Apurva Badithela (Pasadena, CA); Tung Phan (Garden Grove, CA)
Assignee: Motional AD LLC
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,562
App. No.
17/967,725
Granted
May 2, 2023
Kind
B1
Abstract

Provided are methods for counterexample-guided update of a motion planner, which may include training a machine learning model to generate a trajectory for a vehicle. The training may generate a first trained machine learning model having a first plurality of feature weights. The machine learning model may be trained based on a counterexample for which the first trained machine learning model fails to generate a correct trajectory for the vehicle. The training may generate a second trained machine learning model having a second plurality of feature weights. A third plurality of feature weights may be determined by updating the first plurality of feature weights based on the second plurality of feature weights. The machine learning model may be updated to generate the trajectory of the vehicle by applying the third plurality of feature weights. Systems and computer program products are also provided.

Claims (62)

1. A method, comprising:

training, using at least one processor, a machine learning model to generate a trajectory for a vehicle, the training generating a first trained machine learning model having a first plurality of feature weights;

training, using the at least one processor and based on a first counterexample for which the first trained machine learning model fails to generate a correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a second trained machine learning model having a second plurality of feature weights;

determining, using the at least one processor, a third plurality of feature weights by at least updating the first plurality of feature weights based on the second plurality of feature weights; and

updating, using the at least one processor, the machine learning model to generate the trajectory of the vehicle by applying the third plurality of feature weights.

2. The method of claim 1 , further comprising:

training, using the at least one processor and based on a second counterexample for which the first trained machine learning model fails to generate the correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a third trained machine learning model having a fourth plurality of feature weights; and

determining, using the at least one processor, the third plurality of feature weights by at least updating the first plurality of feature weights based on the fourth plurality of feature weights.

3. The method of claim 1 , wherein the updating of the first plurality of feature weights includes updating a probability distribution of the first plurality of feature weights based on the second plurality of the feature weights.

4. The method of claim 3 , wherein the probability distribution of the first plurality of feature weights comprises a Gaussian distribution.

5. The method of claim 3 , wherein the probability distribution of the first plurality of feature weights is updated by performing a statistical inference based on the second plurality of feature weights.

6. The method of claim 5 , wherein the statistical inference comprises a Bayesian inference.

7. The method of claim 1 , wherein the machine learning model comprises a neural network.

8. The method of claim 1 , wherein the machine learning model is trained through inverse reinforcement learning (IRL).

9. The method of claim 1 , further comprising:

selecting, using the at least one processor, a scenario including the vehicle and one or more objects present in an environment of the vehicle, the scenario being selected based on a probability of the first trained machine learning model failing to generate the correct trajectory for the scenario; and

generating, using the at least one processor and based on the scenario, the first counterexample.

10. The method of claim 9 , wherein the scenario is selected from a plurality of scenarios by performing one or more of importance sampling, rejection sampling, Markov Chain Monte Carlo (MCMC), Metropolis-Hastings, Gibbs sampling, slice sampling, or exact sampling.

11. A system, comprising:

at least one processor; and

at least one memory storing instructions thereon that, when executed by the at least one processor, result in operations comprising:

training, using at least one processor, a machine learning model to generate a trajectory for a vehicle, the training generating a first trained machine learning model having a first plurality of feature weights;

training, using the at least one processor and based on a first counterexample for which the first trained machine learning model fails to generate a correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a second trained machine learning model having a second plurality of feature weights;

determining, using the at least one processor, a third plurality of feature weights by at least updating the first plurality of feature weights based on the second plurality of feature weights; and

updating, using the at least one processor, the machine learning model to generate the trajectory of the vehicle by applying the third plurality of feature weights.

12. The system of claim 11 , wherein the operations further comprise:

training, using the at least one processor and based on a second counterexample for which the first trained machine learning model fails to generate the correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a third trained machine learning model having a fourth plurality of feature weights; and

determining, using the at least one processor, the third plurality of feature weights by at least updating the first plurality of feature weights based on the fourth plurality of feature weights.

13. The system of claim 11 , wherein the updating of the first plurality of feature weights includes updating a probability distribution of the first plurality of feature weights based on the second plurality of the feature weights.

14. The system of claim 13 , wherein the probability distribution of the first plurality of feature weights comprises a Gaussian distribution.

15. The system of claim 11 , wherein the probability distribution of the first plurality of feature weights is updated by performing a statistical inference based on the second plurality of feature weights.

16. The system of claim 15 , wherein the statistical inference comprises a Bayesian inference.

17. The system of claim 11 , wherein the machine learning model comprises a neural network.

18. The system of claim 11 , wherein the machine learning model is trained through inverse reinforcement learning (IRL).

19. The system of claim 11 , wherein the operations further comprise:

selecting, using the at least one processor, a scenario including the vehicle and one or more objects present in an environment of the vehicle, the scenario being selected based on a probability of the first trained machine learning model failing to generate the correct trajectory for the scenario; and

generating, using the at least one processor and based on the scenario, the first counterexample.

20. The system of claim 19 , wherein the scenario is selected from a plurality of scenarios by performing one or more of importance sampling, rejection sampling, Markov Chain Monte Carlo (MCMC), Metropolis-Hastings, Gibbs sampling, slice sampling, or exact sampling.

21. A non-transitory storage media storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

training, using at least one processor, a machine learning model to generate a trajectory for a vehicle, the training generating a first trained machine learning model having a first plurality of feature weights;

training, using the at least one processor and based on a first counterexample for which the first trained machine learning model fails to generate a correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a second trained machine learning model having a second plurality of feature weights;

determining, using the at least one processor, a third plurality of feature weights by at least updating the first plurality of feature weights based on the second plurality of feature weights; and

updating, using the at least one processor, the machine learning model to generate the trajectory of the vehicle by applying the third plurality of feature weights.

22. A method, comprising:

applying, using at least one data processor, a machine learning model trained to generate a trajectory for a vehicle, the training of the machine learning model includes

generating a first trained machine learning model having a first plurality of feature weights,

training, using the at least one processor and based on a first counterexample for which the first trained machine learning model fails to generate a correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a second trained machine learning model having a second plurality of feature weights,

determining, using the at least one processor, a third plurality of feature weights by at least updating the first plurality of feature weights based on the second plurality of feature weights, and

updating, using the at least one processor, the machine learning model to generate the trajectory of the vehicle by applying the third plurality of feature weights; and

controlling, using the at least one processor, a motion of the vehicle based at least on the trajectory.

23. The method of claim 22 , wherein the training of the machine learning model further includes

training, using the at least one processor and based on a second counterexample for which the first trained machine learning model fails to generate the correct trajectory for the vehicle, the machine learning model to generate the correct trajectory, the training generating a third trained machine learning model having a fourth plurality of feature weights, and

determining, using the at least one processor, the third plurality of feature weights by at least updating the first plurality of feature weights based on the fourth plurality of feature weights.

24. The method of claim 22 , wherein the updating of the first plurality of feature weights includes updating a probability distribution of the first plurality of feature weights based on the second plurality of the feature weights.

25. The method of claim 24 , wherein the probability distribution of the first plurality of feature weights comprises a Gaussian distribution.

26. The method of claim 24 , wherein the probability distribution of the first plurality of feature weights is updated by performing a statistical inference based on the second plurality of feature weights.

27. The method of claim 26 , wherein the statistical inference comprises a Bayesian inference.

28. The method of claim 22 , wherein the machine learning model comprises a neural network.

29. The method of claim 22 , wherein the machine learning model is trained through inverse reinforcement learning (IRL).

30. The method of claim 22 , further comprising:

selecting, using the at least one processor, a scenario including the vehicle and one or more objects present in an environment of the vehicle, the scenario being selected based on a probability of the first trained machine learning model failing to generate the correct trajectory for the scenario; and

generating, using the at least one processor and based on the scenario, the first counterexample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: BADITHELA, APURVA; PHAN, TUNG
To: MOTIONAL AD LLC
Reel/Frame 061676/0919 →
Continuity (1)
Provisional Application 63303202 · Jan 26, 2022
Cited By (4)
US 12,194,631 US 12,371,025 US 12,565,234 US 12,623,685