IP Library › Granted Patent US 12,280,499
Granted Patent B2
US 12,280,499 · App. 17/095,586 · Granted Apr 22, 2025

Domain adaptation for simulated motor backlash

Inventors: Sergey Bashkirov (Salida, CA); Michael Taylor (San Mateo, CA)
Assignee: Sony Interactive Entertainment Inc.
B25J9/163B25J9/12G06F30/27G06N3/084G06V10/751G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,280,499
App. No.
17/095,586
Filed
Nov 11, 2020
Granted
Apr 22, 2025
Kind
B2
Examiner
JEONG, HEIN
Art Unit
2188
USPC
703/7
Abstract

A method, system and computer program product for training a control input system involve taking an integral of an output value from a Motion Decision Neural Network for one or more movable joints to generate an integrated output value and comparing the integrated output value to a backlash threshold. A subsequent output value is generated using a machine learning algorithm that includes a sensor value and a previous joint position if the integrated output value does not at least meet the threshold. A position of the one or more movable joints is simulated based on an integral of the subsequent output value; and the Motion Decision Neural Network is trained with the machine learning algorithm based upon at least a result of the simulation of the position of the one or more movable joints.

Claims (44)

1. A method for training a control input system, comprising:

a) taking an integral of an output value from a Motion Decision Neural Network for one or more movable joints to generate an integrated output value;

b) comparing the integrated output value to a backlash threshold;

c) generating a subsequent output value using a machine learning algorithm that includes as inputs a sensor value and a previous joint position when the integrated output value does not at least meet the threshold;

d) simulating a position of the one or more movable joints based on an integral of the subsequent output value; and

e) training the Motion Decision Neural Network with the machine learning algorithm based upon at least a result of the simulation of the position of the one or more movable joints;

repeating a) through e), wherein c) includes generating the subsequent output value using a machine learning algorithm that includes a sensor value and the integrated output value when the integrated output value meets or exceeds the threshold; and

controlling a robot by passing the integrated output value to a movable joint when the integrated output value meets or exceeds the threshold.

2. The method of claim 1 , wherein the movable joint is in a virtual simulation.

3. The method of claim 1 , wherein the movable joint is a motorized joint.

4. The method of claim 1 , wherein the integral of the output value is a first integral of the output value and the integral of the subsequent output value is a first integral of the subsequent output value.

5. The method of claim 1 , wherein the integral of the output value is a second integral of the output value and the integral of the subsequent output value is a second integral of the subsequent output value.

6. The method of claim 5 , further comprising taking a first integral of the output value and wherein the integrated output value includes the first integral of the output value and the second integral of the output value.

7. The method of claim 1 , wherein the threshold is a randomized value.

8. The method of claim 7 , wherein steps a) through e) are repeated with a different randomized value each repetition.

9. The method of claim 1 , wherein the threshold is at least based on a time or a number of repetitions of steps a) through e).

10. The method of claim 9 , wherein the threshold value changes with a change in the time.

11. The method of claim 10 , wherein at least one of the one or more movable joints is associated with a threshold value that changes differently with the change in time than another movable joint of the one or more movable joints.

12. The method of claim 1 , wherein the threshold is different for at least one of the one or more movable joint compared to another joint of the one or more movable joints.

13. The method of claim 12 , wherein the different threshold depends at least upon a location of the at least one of the one or more movable joint in a simulation.

14. The method of claim 1 , wherein the threshold value at least depends upon an angle value of at least one of the one or more movable joints.

15. The method of claim 14 , wherein the threshold value changes based on the amount of times the at least one of the one or more joints passes through the angle value.

16. The method of claim 14 wherein the angle dependent threshold values are represented by randomized heat map to simulate wear in the at least one of the one or more different joints or transitions between surfaces.

17. The method of claim 16 wherein the heatmap simulates a transition of the joints from areas of high use to areas of lower use.

18. The method of claim 17 , wherein different angles of the joints have different threshold valued for different surface types.

19. A input control system comprising:

a processor;

a memory coupled to the processor;

non-transitory instruction embedded in the memory that when executed by the processor cause the processor to carry out the method for training control input comprising:

a) taking an integral of an output value from a Motion Decision Neural Network for one or more movable joints to generate an integrated output value;

b) comparing the integrated output value to a backlash threshold;

c) generating a subsequent output value using a machine learning algorithm that includes as inputs a simulated sensor value and a previous simulated joint position when the integrated output value does not at least meet the threshold;

d) simulating a position of the one or more simulated movable joints based on an integral of the subsequent output value; and

e) training the Motion Decision Neural Network with the machine learning algorithm based upon at least a result of the simulation of the position of the one or more simulated movable joints;

repeating a) through e), wherein c) includes generating the subsequent output value using a machine learning algorithm that includes a sensor value and the integrated output value when the integrated output value meets or exceeds the threshold; and

controlling a robot by passing the integrated output value to a movable joint when the integrated output value meets or exceeds the threshold.

20. A non-transitory computer readable medium having instruction embedded thereon that when executed cause a computer to carry out the method for training a control input system comprising:

a) taking an integral of an output value from a Motion Decision Neural Network for one or more movable joints to generate an integrated output value;

b) comparing the integrated output value to a backlash threshold;

c) generating a subsequent output value using a machine learning algorithm that includes as inputs a sensor value and a previous joint position when the integrated output value does not at least meet the threshold;

d) simulating a position of the one or more movable joints based on an integral of the subsequent output value; and

e) training the Motion Decision Neural Network with the machine learning algorithm based upon at least a result of the simulation of the position of the one or more movable joints;

repeating a) through e), wherein c) includes generating the subsequent output value using a machine learning algorithm that includes a sensor value and the integrated output value when the integrated output value meets or exceeds the threshold; and

controlling a robot by passing the integrated output value to a movable joint when the integrated output value meets or exceeds the threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2020
From: TAYLOR, MICHAEL; BASHKIROV, SERGEY
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 054341/0481 →
Continuity (1)
Related Publication 20220143820A1 · May 12, 2022
References Cited (72)
US 5719480A · Bock et al. · 1998 [cited by applicant]
US 9786202B2 · Huang et al. · 2017 [cited by applicant]
US 10937240B2 · Anderson · 2021 [cited by applicant]
US 11278433B2 · Herr et al. · 2022 [cited by applicant]
US 20030216895A1 · Ghaboussi et al. · 2003 [cited by applicant]
US 20050179417A1 · Takenaka · 2005 [cited by examiner]
US 20070255454A1 · Dariush · 2007 [cited by applicant]
US 20080027582A1 · Obinata et al. · 2008 [cited by applicant]
US 20110270443A1 · Kamiya et al. · 2011 [cited by applicant]
US 20120166165A1 · Nogami et al. · 2012 [cited by applicant]
US 20140107841A1 · Danko · 2014 [cited by applicant]
US 20160361633A1 · Fujita et al. · 2016 [cited by applicant]
US 20170083047A1 · Helot et al. · 2017 [cited by applicant]
US 20170183047A1 · Takagi · 2017 [cited by examiner]
US 20170252922A1 · Levine et al. · 2017 [cited by applicant]
US 20170334066A1 · Levine et al. · 2017 [cited by applicant]
US 20180001206A1 · Osman et al. · 2018 [cited by applicant]
US 20180250812A1 · Narukawa · 2018 [cited by applicant]
US 20190324538A1 · Rihn et al. · 2019 [cited by applicant]
US 20190384316A1 · Qi et al. · 2019 [cited by applicant]
US 20200139539A1 · Hasunuma et al. · 2020 [cited by applicant]
US 20200143229A1 · Made et al. · 2020 [cited by applicant]
US 20200160535A1 · Akbarian et al. · 2020 [cited by applicant]
US 20200218365A1 · Todorov et al. · 2020 [cited by applicant]
US 20200290203A1 · Taylor et al. · 2020 [cited by applicant]
US 20200293881A1 · Taylor · 2020 [cited by applicant]
US 20200319721A1 · Erivantcev et al. · 2020 [cited by applicant]
US 20210158141A1 · Taylor et al. · 2021 [cited by applicant]
US 20220143821A1 · Bashkirov et al. · 2022 [cited by applicant]
US 20220143822A1 · Bashkirov et al. · 2022 [cited by applicant]
CA 3133873A1 · 2020 [cited by examiner]
CN 107610208A · 2018 [cited by applicant]
CN 110007750A · 2019 [cited by applicant]
CN 111653010A · 2020 [cited by examiner]
CN 111898571A · 2020 [cited by applicant]
EP 1559460A1 · 2005 [cited by applicant]
EP 2871575A1 · 2015 [cited by applicant]
TW 201905729A · 2019 [cited by applicant]
TW 202026846A · 2020 [cited by applicant]
WO 2012079541A1 · 2012 [cited by applicant]
WO 2017129200A1 · 2017 [cited by applicant]
WO 2020008878A1 · 2020 [cited by applicant]
Kiemel, J. C. et al. TrueÆdapt: Learning smooth online trajectory adaptation with bounded jerk, acceleration and velocity in joint space. Oct. 2020. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Sy… [cited by examiner]
Kiemel, J. C. et al. Learning Robot Trajectories subject to Kinematic Joint Constraints. Nov. 1, 2020. arXiv preprint arXiv: 2011.00563. (Year: 2020). [cited by examiner]
Bazzi, D., Messeri, C., Zanchettin, A. M., & Rocco, P. Identification of robot forward dynamics via neural network. Oct. 2020. In 2020 4th International Conference on Automation, Control and Robots (ICACR) (pp. 13-21). … [cited by examiner]
Non-Final Office Action for U.S. Appl. No. 16/693,093, dated Sep. 27, 2023. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/095,617, dated Oct. 26, 2023. [cited by applicant]
Non-Final/Final Office Action for U.S. Appl. No. 17/095,640, dated Sep. 13, 2023. [cited by applicant]
Arthur Juliani “Simple Reinforcement Learning with Tensorflow Part 8: Asynchronous Actor-Critic Agents (A3C)” Medium, Dated: Dec. 16, 2016, Available at: https://medium.com/emergent-future/simple-reinforcement-learning-… [cited by applicant]
Arthur Juliani, “Simple reinforcement Learning with Tensorflow Part 0: Q-learning with Tables and Neural Networks” Medium, dated: Aug. 25, 2016. Available at: https://medium.com/emergent-future/simple-reinforcement-lear… [cited by applicant]
Arthur Juliani, “Simple Reinforcement Learning with Tensorflow: Part 2—Policy-based Agents” Medium, Dated: Jun. 24, 2016. Available at: https://medium.com/@awjuliani/super-simple-reinforcement-learning-tutorial-part-2-d… [cited by applicant]
Coulom, Remi “Apprentissage par renforcement utilsant des resaux de neurones, avec des applications au control moteur” PHD Thesis Institute Nat. Polytech. of Grenoble Jun. 19, 2002. [cited by applicant]
Ghoshavi, Abhijit “Neural Networks and Reinforcement Learning” Slide presentation Dept. of Eng. Mgmt. and Sys. Eng. Missouri Univ. of Sci. and Tech. Accessed Nov. 6, 2019. [cited by applicant]
Kniewasser, Gerhard et al. “Reinforcemetn Learning with Dynamic Movement Primitives—DMPs” Dept. of Theoretical Comp. Sci. Univ. of Tech., Graz (2013) Accessed Nov. 6, 2019. [cited by applicant]
Kober, Jens et al. “Reinforcement Learning to Adjust Robot Movements to New Situations” Proceedings of the Twenty-Second Int. Joint Conf. on Artificial Intelligence p. 2650-2655 Accessed Nov. 6, 2019. [cited by applicant]
Peng, Xue Bin, and Michiel van de Panne. “Learning Locomotion Skills Using DeepRL.” Proceedings of the ACM SIGGRAPH / Eurographics Symposium on Computer Animation—SCA '17 (2017). [cited by applicant]
Sepp Hochreiter et al., “Long Short-Term Memory”, Neural Computation, 9(8) 1735-1780, 1997, Cambridge, MA. [cited by applicant]
U.S. Appl. No. 16/693,093 to Michael Taylor and Sergey Bashkirov filed Nov. 22, 2019. [cited by applicant]
Zaiqiang Wu, et al. “Analytical Derivatives for Differentiable Renderer: 3D Pose estimation by Silhouette Consistency” Zhejiang University, ArXiv:1906.07870v1, Jun. 19, 2019 available at: https://arxiv.org/pdf/1906.0787… [cited by applicant]
Erlhagen et al, “The dynamic neural field approach to cognitive robotics” In J. Neural Eng. 3 (2006), R36-R54, [online] [retrieved on Jan. 5, 2022 (Jan. 5, 2022] Retrieved from the Internet. [cited by applicant]
Hu et al, “Impedance with Finite-Time Control Scheme for Robot-Environment Interaction” in Mathematical Problems in Engineering, vol. 2020, Article ID 2796590 , 18 Pages May 25, 2020, [online] [retrieved on Jan. 5, 2022… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2021/58390, Feb. 7, 2022. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2021/58392, Feb. 7, 2022. [cited by applicant]
Peng et al, “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization” in arXiv:1710.06537, 8 Pages, [Submitted on Oct. 18, 2017 (v1), last revised Mar. 3, 2018 (this version, v3)], [online] [retrieved on Apr… [cited by applicant]
Taiwanese Office Action for Country Code Application No. 893689, dated Jun. 28, 2022. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2021/58388, Feb. 4, 2022. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/693,093, dated Jan. 19, 2024. [cited by applicant]
Final Office Action for U.S. Appl. No. 17/095,617, dated Apr. 19, 2024. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/693,093, dated Jun. 11, 2024. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/095,640, dated Feb. 22, 2024. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/693,093, dated Aug. 15, 2024. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/095,617, dated Dec. 4, 2024. [cited by applicant]