IP Library › Granted Patent US 12,474,704
Granted Patent B2
US 12,474,704 · App. 17/774,604 · Granted Nov 18, 2025

Robot control model learning method for reducing frequency of robot intervention behavior

Inventors: Mai Kurose (Tokyo, JP); Ryo Yonetani (Tokyo, JP)
Assignee: OMRON CORPORATION
G05D1/0088G05D1/0214
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,474,704
App. No.
17/774,604
Granted
Nov 18, 2025
Kind
B2
Abstract

A robot control model learning device ( 10 ) performs, by using state information indicating the state of a robot which autonomously travels to a destination in a dynamic environment as an input, reinforcement learning to obtain a robot control model for selecting and outputting a behavior in accordance with the state of the robot from among a plurality of behaviors including an intervention behavior for intervening in the environment, while using the number of times the intervention behavior has been performed as a minus reward.

Claims (88)

1 . A robot control model learning method, comprising, by a computer:

subjecting a robot control model to reinforcement learning, the robot control model being input with state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment, and the robot control model selecting and outputting a behavior corresponding to the state of the robot from among a plurality of behaviors, wherein:

the plurality of behaviors include an intervention behavior of intervening in the dynamic environment,

the intervention behavior is a non-moving behavior of the robot to notify someone of an existence of the robot in order for the robot to travel to the destination without stopping,

the reinforcement learning is performed on the robot control model using a reward function,

the reward function includes a weighted combination of an environment reward and an influence reward,

the environment reward takes a first negative value when the robot collides with another object,

the influence reward takes a second negative value when the robot executes the intervention behavior, and

the reinforcement learning comprises:

learning a deep neural network representing a state value function,

determining the state of the robot based on: the reward function, a discount factor, an increase in time in one step, and the state value function, and

updating the state value function by updating parameters of the deep neural network.

2 . The robot control model learning method of claim 1 , wherein:

the plurality of behaviors include at least one of a moving direction of the robot, a moving velocity of the robot, or the intervention behavior, and

the reward function is given such that at least one of an arrival time until the robot reaches the destination or a number of times of intervention decreases.

3 . The robot control model learning method of claim 1 , wherein:

the plurality of behaviors include an avoidance behavior, which is a moving behavior of the robot to avoid collision with another object when travelling to the destination, and

the reward function is given such that a number of times of avoidance in which the collision is avoided decreases.

4 . The robot control model learning method of claim 1 , wherein:

the discount factor is set to a value between zero and one, and

a reward that is obtained at a later time is to be discounted more than a reward that is obtained at an earlier time, based on the discount factor and the increase in time in one step.

5 . A robot control model learning device, comprising:

a learning section that includes a processor and a memory and subjects a robot control model to reinforcement learning, the robot control model being configured to receive state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment, and to select and output a behavior corresponding to the state of the robot from among a plurality of behaviors, wherein:

the plurality of behaviors include an intervention behavior of intervening in the dynamic environment,

the intervention behavior is a non-moving behavior of the robot to notify someone of an existence of the robot in order for the robot to travel to the destination without stopping,

the reinforcement learning is performed on the robot control model using a reward function,

the reward function includes a weighted combination of an environment reward and an influence reward,

the environment reward takes a first negative value when the robot collides with another object,

the influence reward takes a second negative value when the robot executes the intervention behavior, and

the reinforcement learning comprises:

learning a deep neural network representing a state value function,

determining the state of the robot based on: the reward function, a discount factor, an increase in time in one step, and the state value function, and

updating the state value function by updating parameters of the deep neural network.

6 . A non-transitory recording medium storing a robot control model learning program that is executable by a computer to perform processing, the processing comprising:

subjecting a robot control model to reinforcement learning, the robot control model being input with state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment, and the robot control model selecting and outputting a behavior corresponding to the state of the robot from among a plurality of behaviors, wherein:

the plurality of behaviors include an intervention behavior of intervening in the dynamic environment,

the intervention behavior is a non-moving behavior of the robot to notify someone of an existence of the robot in order for the robot to travel to the destination without stopping,

the reinforcement learning is performed on the robot control model using a reward function,

the reward function includes a weighted combination of an environment reward and an influence reward,

the environment reward takes a first negative value when the robot collides with another object,

the influence reward takes a second negative value when the robot executes the intervention behavior, and

the reinforcement learning comprises:

learning a deep neural network representing a state value function,

determining the state of the robot based on: the reward function, a discount factor, an increase in time in one step, and the state value function, and

updating the state value function by updating parameters of the deep neural network.

7 . A robot control method, according to which a computer performs processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 1 .

8 . A robot control device, comprising:

an acquisition section that is implemented on a processor and a memory and acquires state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

a control section that is implemented on a processor and a memory and effects control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning device of claim 5 .

9 . A non-transitory recording medium storing a robot control program that is executable by a computer to perform processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 1 .

10 . A robot, comprising:

an acquisition section that is implemented on a processor and a memory and acquires state information expressing a state of the robot, which travels autonomously to a destination in a dynamic environment;

an autonomous traveling section that includes a motor and causes the robot to travel autonomously; and

a robot control device that includes a control section implemented on a processor and a memory and configured to effect control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning device of claim 5 .

11 . A robot control method, according to which a computer performs processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 2 .

12 . A robot control method, according to which a computer performs processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 3 .

13 . A robot control method, according to which a computer performs processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 1 .

14 . A non-transitory recording medium storing a robot control program that is executable by a computer to perform processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 2 .

15 . A non-transitory recording medium storing a robot control program that is executable by a computer to perform processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 3 .

16 . A non-transitory recording medium storing a robot control program that is executable by a computer to perform processing, the processing comprising:

acquiring state information expressing a state of a robot that travels autonomously to a destination in a dynamic environment; and

effecting control such that the robot moves to the destination on the basis of the state information and on the basis of the robot control model learned by the robot control model learning method of claim 1 .

17 . The robot control model learning device of claim 5 , wherein:

the plurality of behaviors include at least one of a moving direction of the robot, a moving velocity of the robot, or the intervention behavior, and

the reward function is given such that at least one of an arrival time until the robot reaches the destination or a number of times of intervention decreases.

18 . The robot control model learning method of claim 5 , wherein:

the plurality of behaviors include an avoidance behavior, which is a moving behavior of the robot to avoid a collision with another object when travelling to the destination, and

the reward function is given such that a number of times of avoidance in which the collision is avoided decreases.

19 . The non-transitory recording medium of claim 6 , wherein:

the plurality of behaviors include at least one of a moving direction of the robot, a moving velocity of the robot, or the intervention behavior, and

the reward function is given such that at least one of an arrival time until the robot reaches the destination or a number of times of intervention decreases.

20 . The non-transitory recording medium of claim 6 , wherein:

the plurality of behaviors include an avoidance behavior, which is a moving behavior of the robot to avoid a collision with another object when travelling to the destination, and

the reward function is given such that a number of times of avoidance in which the collision is avoided decreases.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2022
From: KUROSE, MAI; YONETANI, RYO
To: OMRON CORPORATION
Reel/Frame 059835/0693 →
Priority Claims (1)
JP 2019-205690 · Nov 13, 2019 · national
Continuity (1)
Related Publication 20220397900A1 · Dec 15, 2022
References Cited (31)
US 10872379B1 · Nepomuceno · 2020 [cited by examiner]
US 11199853B1 · Afrouzi · 2021 [cited by examiner]
US 20100194593A1 · Mays · 2010 [cited by examiner]
US 20110144850A1 · Jikihara · 2011 [cited by applicant]
US 20130184980A1 · Ichikawa · 2013 [cited by examiner]
US 20190072959A1 · Palanisamy · 2019 [cited by examiner]
US 20190272558A1 · Suzuki et al. · 2019 [cited by applicant]
US 20190310632A1 · Nakhaei Sarvedani · 2019 [cited by examiner]
US 20200069134A1 · Ebrahimi Afrouzi · 2020 [cited by examiner]
US 20200249674A1 · Dally · 2020 [cited by examiner]
US 20210171025A1 · Ishikawa · 2021 [cited by examiner]
CN 103901887A · 2014 [cited by applicant]
CN 107440891A · 2017 [cited by applicant]
CN 109213148A · 2019 [cited by applicant]
CN 109828574A · 2019 [cited by applicant]
JP 2012139798A · 2012 [cited by applicant]
JP 2012187698A · 2012 [cited by applicant]
JP 2018198012A · 2018 [cited by applicant]
JP 2019096012A · 2019 [cited by applicant]
WO 2012039280A1 · 2012 [cited by applicant]
WO 2018110305A1 · 2018 [cited by applicant]
Sasaki et al., A3C Based Motion Learning for an Autonomous Mobile Robot in Crowds (Year: 2019). [cited by examiner]
Extended European search report issued in corresponding International Application No. 20888239.9 dated Nov. 13, 2023. [cited by applicant]
Sasaki et al., “A3C Based Motion Learning for an Autonomous Mobile Robot in Crowds”, 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC), IEEE, Oct. 6-9, 2019, pp. 1036-1042, XP033667706. [cited by applicant]
Wang et al., “Improved Multi-Agent Reinforcement Learning for Path Planning-Based Crowd Simulation”, IEEE Access, vol. 7, Jun. 19, 2019, pp. 73841-73855, XP011731208. [cited by applicant]
Chen et al., “Decentralized Non-communicating Multiagent Collision Avoidance with Deep Reinforcement Learning,” https://arxiv.org/pdf/1609.07845 (2016). [cited by applicant]
Chen et al., “Socially Aware Motion Planning with Deep Reinforcement Learning,” https://arxiv.org/pdf/1703.08862.pdf (2018). [cited by applicant]
ZMP https://news.mynavi.jp/article/20180323-604926/ (Mar. 23, 2018). [cited by applicant]
International Search Report issued in corresponding International Application No. PCT/JP2020/039554 dated Dec. 8, 2020. [cited by applicant]
Written Opinion of the International Searching Authority issued in corresponding International Application No. PCT/JP2020/039554 dated Dec. 8, 2020. [cited by applicant]
Office Action issued in corresponding Chinese Patent Application No. 202080076868.3, dated Nov. 12, 2024, with a partial English translation. [cited by applicant]