IP Library › Granted Patent US 12,194,637
Granted Patent B2
US 12,194,637 · App. 17/963,017 · Granted Jan 14, 2025

Imitation learning and model integrated trajectory planning

Inventors: Zhixian Ye (Santa Clara, CA); Qiangqiang Guo (Seattle, WA); Liyang Wang (Sunnyvale, CA); Liangjun Zhang (Cupertino, CA)
Assignee: Baidu USA LLC
B25J9/1664B25J9/163
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,194,637
App. No.
17/963,017
Granted
Jan 14, 2025
Kind
B2
Abstract

Presented herein are embodiments of a two-stage methodology that integrates data-driven imitation learning and model-based trajectory optimization to generate optimal trajectories for autonomous excavators. In one or more embodiments, a deep neural network using demonstration data to mimic the operation patterns of human experts under various terrain states, including their geometry shape and material type. A stochastic trajectory optimization methodology is used to improve the trajectory generated by the neural network to ensure kinematics feasibility, improve smoothness, satisfy hard constraints, and achieve desired excavation volumes. Embodiments were tested on a Franka robot arm equipped with a bucket end-effector. Embodiments were also evaluated on different material types, such as sand and rigid blocks. Experimental results showed that embodiments of the two-stage methodology that comprises combining expert knowledge and model optimization increased the excavation weights by up to 24.77% with low variance.

Claims (55)

1. A computer-implemented method for training a system to generate a trajectory, the method comprising:

using an actor neural network, which receives a state as an input, to obtain an initial trajectory shape;

using the initial trajectory shape as an input into a trajectory motion planning model to obtain an operational trajectory;

checking whether the initial trajectory shape is acceptable;

responsive to the initial trajectory shape not being acceptable, having a human generate a human-demonstrated trajectory for the state;

having a robot execute an execution trajectory that is either the operational trajectory or the human-demonstrated trajectory;

after execution, evaluating the execution trajectory and assigning a value to it;

storing a set of data comprising the state, the execution trajectory, and the value as new demonstration data; and

updating the actor neural network using an aggregated demonstration data, wherein the actor neural network is initially pre-trained using an original set of demonstration data, and the aggregated demonstration data set comprises the new demonstration data and the original set of demonstration data.

2. The computer-implemented method of claim 1 wherein the trajectory motion planning model comprises:

a stochastic trajectory optimization for motion planning (STOMP) optimization model.

3. The computer-implemented method of claim 2 wherein the STOMP optimization model considers the robot's one or more kinematic models to consider feasibility of the operational trajectory and further comprises one or more:

one or more hard constraints;

a smoothness of the operational trajectory; and

one or more user-defined objectives.

4. The computer-implemented method of claim 2 wherein the STOMP optimization model is trained using a trajectory-based cost function.

5. The computer-implemented method of claim 1 further comprising:

repeating the steps of claim 1 , in which a new state is obtain for each iteration, until a stop condition is reached; and

outputting a trained actor neural network.

6. The computer-implemented method of claim 1 wherein the step of updating the actor neural network comprises:

training the actor neural network to minimize a mean square error (MSE) between demonstration trajectories and predicted trajectories.

7. The computer-implemented method of claim 1 wherein the state comprises a set of one or more features obtained from a terrain point cloud that relate to geometric shape of a terrain and a set of one or more features obtained from an image classifier that relate to material composition of the terrain.

8. The computer-implemented method of claim 1 wherein the execution trajectory is for excavation.

9. A system comprising:

one or more processors; and

a non-transitory computer-readable medium or media comprising one or more sets of instructions which, when executed by at least one of the one or more processors, causes steps to be performed comprising:

using an actor neural network, which receives a state as an input, to obtain an initial trajectory shape;

using the initial trajectory shape as an input into a trajectory motion planning model to obtain an operational trajectory;

checking whether the initial trajectory shape is acceptable;

responsive to the initial trajectory shape not being acceptable, having a human generate a human-demonstrated trajectory for the state;

having a robot execute an execution trajectory that is either the operational trajectory or the human-demonstrated trajectory;

after execution, evaluating the execution trajectory and assigning a value to it;

storing a set of data comprising the state, the execution trajectory, and the value as new demonstration data; and

updating the actor neural network using an aggregated demonstration data set, wherein the actor neural network is initially pre-trained using an original set of demonstration data, and the aggregated demonstration data set comprises the new demonstration data and the original set of demonstration data.

10. The system of claim 9 wherein the trajectory motion planning model comprises:

a stochastic trajectory optimization for motion planning (STOMP) optimization model.

11. The system of claim 10 wherein the STOMP optimization model considers the robot's one or more kinematic models to consider feasibility of the operational trajectory and further comprises one or more:

one or more hard constraints;

a smoothness of the operational trajectory; and

one or more user-defined objectives.

12. The system of claim 10 wherein the STOMP optimization model is trained using a trajectory-based cost function.

13. The system of claim 9 wherein the non-transitory computer-readable medium or media further comprises one or more sequences of instructions which, when executed by at least one of the one or more processors, causes steps to be performed comprising:

repeating the steps of claim 9 , in which a new state is obtain for each iteration, until a stop condition is reached; and

outputting a trained actor neural network.

14. The system of claim 9 wherein the step of updating the actor neural network comprises:

training the actor neural network to minimize a mean square error (MSE) between demonstration trajectories and predicted trajectories.

15. The system of claim 9 wherein the state comprises a set of one or more features obtained from a terrain point cloud that relate to geometric shape of a terrain and a set of one or more features obtained from an image classifier that relate to material composition of the terrain.

16. A computer-implemented method for generating a trajectory, the method comprising:

inputting a state into a two-stage trajectory generating system comprising a trained actor neural network that was trained using imitation learning and a model-based trajectory motion planner;

using the trained actor neural network, which receives the state as an input, to obtain an initial trajectory shape;

using the initial trajectory shape as an input into a model-based trajectory motion planner to obtain an execution trajectory; and

providing the execution trajectory for use with a device that is to execute the execution trajectory.

17. The computer-implemented method of claim 16 wherein the model-based trajectory motion planner comprises:

a stochastic trajectory optimization for motion planning (STOMP) optimization model.

18. The computer-implemented method of claim 16 wherein the state comprises a set of one or more features obtained from a terrain point cloud that relate to geometric shape of a terrain and a set of one or more features obtained from an image classifier that relate to material composition of the terrain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: YE, ZHIXIAN; GUO, QIANGQIANG; WANG, LIYANG; ZHANG, LIANGJUN
To: BAIDU USA, LLC
Reel/Frame 062172/0917 →
Continuity (1)
Related Publication 20240131705A1 · Apr 25, 2024
References Cited (23)
US 11409287B2 · Zhang · 2022 [cited by examiner]
US 12017352B2 · Chadalavada Vijay Kumar · 2024 [cited by examiner]
US 20210086364A1 · Handa · 2021 [cited by examiner]
US 20210223774A1 · Zhang · 2021 [cited by examiner]
US 20210292998A1 · Kawamoto · 2021 [cited by examiner]
US 20220126445A1 · Zhu · 2022 [cited by examiner]
US 20220134537A1 · Chadalavada Vijay Kumar · 2022 [cited by examiner]
US 20230036849A1 · Lu · 2023 [cited by examiner]
CN 114723020A · 2022 [cited by examiner]
CN 115330055A · 2022 [cited by examiner]
EP 4083335A2 · 2022 [cited by examiner]
WO WO2019142229A1 · 2019 [cited by examiner]
Task-unit based trajectory generation for excavators utilizing expert operator skills (Year: 2023). [cited by examiner]
B. J. Hodel., “Learning to Operate an Excavator via Policy Optimization,” Procedia Computer Science, vol. 140, 2018, [online], [Retrieved Apr. 25, 2024]. Retrieved from Internet <URL: https://www.sciencedirect.com/scien… [cited by applicant]
I. Kurinov et al.,“Automated excavator based on reinforcement learning and multibody system dynamics,” IEEE Access, vol. 8, 2020. (9pgs). [cited by applicant]
S. Dadhich et al., “Machine learning approach to automatic bucket loading,” in 24th Mediterranean Conference on Control and Automation (MED), 2016. (13pgs). [cited by applicant]
R. Sutton et al., “Reinforcement Learning, second edition: An Introduction,” ser. Adaptive Computation & Machine Learning series, MIT Press, 2018, [online], [Retrieved Apr. 25, 2024]. Retrieved from Internet <URL: https… [cited by applicant]
F. Torabi et al., “Behavioral Cloning from Observation,” Proceedings of the 27th International Joint Conference on Artificial Intelligence, 2018. (8pgs). [cited by applicant]
J. Fu et al., “Learning robust rewards with adversarial inverse reinforcement learning,” arXiv preprint arXiv:1710.11248, 2018. (15pgs). [cited by applicant]
S. Ross et al., “A reduction of imitation learning and structured prediction to no-regret online learning,” arXiv preprint arXiv:1011.0686, 2011. (9pgs). [cited by applicant]
M. Kalakrishnan et al., “STOMP: Stochastic Trajectory Optimization for Motion Planning,” in IEEE International Conference on Robotics & Automation, 2011. (6pgs). [cited by applicant]
L. Zhang et al., “An autonomous excavator system for material loading tasks,” [online], [Retrieved Apr. 24, 2024]. Retrieved from Internet <URL: https://robotics.sciencemag.org/content/6/55/eabc3164> Science Robotics, v… [cited by applicant]
D. Jud et al., “Planning and Control for Autonomous Excavation,” IEEE Robotics & Automation Letters, vol. 2, No. 4, 2017.(8pgs). [cited by applicant]