IP Library › Granted Patent US 11,511,420
Granted Patent B2
US 11,511,420 · App. 15/702,495 · Granted Nov 29, 2022

Machine learning device, robot system, and machine learning method for learning operation program of robot

Inventor: Syuntarou Toda (Yamanashi, JP)
Assignee: FANUC CORPORATION
B25J9/1664B25J9/163B25J19/023G05B13/029G05B13/0265G06N3/008G06N3/08G06N3/084G06N5/022G06N20/00G05B2219/32334G05B2219/33321G05B2219/45104G05B2219/45135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,511,420
App. No.
15/702,495
Granted
Nov 29, 2022
Kind
B2
Abstract

A machine learning device, which learns an operation program of a robot, includes a state observation unit which observes as a state variable at least one of a shaking of an arm of the robot and a length of an operation trajectory of the arm of the robot; a determination data obtaining unit which obtains as determination data a cycle time in which the robot performs processing; and a learning unit which learns the operation program of the robot based on an output of the state observation unit and an output of the determination data obtaining unit.

Claims (37)

1. A machine learning device configured to learn an operation program of a robot, the machine learning device comprising a processor configured to:

observe, as a state variable, a shaking of an arm of the robot and a length of an operation trajectory of the arm of the robot;

obtain, as determination data, a cycle time in which the robot performs processing;

learn the operation program of the robot, based on the state variable and the determination data, by

calculating a reward based on the state variable and the determination data, the reward including positive rewards and negative rewards;

adding the positive rewards and the negative rewards together to obtain an added reward value for each of the shaking of the arm, the length of the operation trajectory of the arm, and the cycle time;

further adding the added reward value for each of the shaking of the arm, the length of the operation trajectory of the arm, and the cycle time together to obtain a further added reward value; and

updating a value function which determines a value of the operation program of the robot based on the state variable, the determination data, and the further added reward value;

determine an operation of the robot based on the learned operation program; and

output the determined operation of the robot to the robot,

wherein the processor is configured to set a negative reward among the negative rewards based on the state variable and the determination data and cause the robot to operate with the shaking of the arm being reduced, and

wherein the processor is configured to set a further negative reward among the negative rewards based on the state variable and the determination data and cause the robot to operate with a reduced cycle time.

2. The machine learning device according to claim 1 , wherein

data on the shaking of the arm and the length of the operation trajectory of the arm is observed based on an image captured by a camera or data from a robot controller, and

the cycle time is obtained from data from the robot controller or by analyzing an image captured by the camera.

3. The machine learning device according to claim 1 , wherein

the processor is further configured to observe, as the state variable, at least one of a position, a speed, and an acceleration of the arm.

4. The machine learning device according to claim 1 , wherein the processor is further configured to determine

an operation of the robot based on the learned operation program.

5. The machine learning device according to claim 1 , wherein the machine learning device further comprises a neural network.

6. The machine learning device according to claim 1 , wherein

the machine learning device is provided in each of a plurality of the robots, at least one of the plurality of robots connectable to at least one other machine learning device, the machine learning device configured to mutually exchange or share a result of machine learning with the at least one other machine learning device.

7. The machine learning device according to claim 1 , wherein

the machine learning device is located in a cloud server or a fog server.

8. A machine learning method for learning an operation program of a robot, the machine learning method comprising:

observing, as a state variable, a shaking of an arm of the robot and a length of an operation trajectory of the arm of the robot;

obtaining, as determination data, a cycle time in which the robot performs processing;

learning the operation program of the robot based on the state variable and the determination data, wherein the learning of the operation program of the robot comprises:

calculating a reward based on the state variable and the determination data, the reward including positive rewards and negative rewards;

adding the positive rewards and the negative rewards together to obtain an added reward value for each of the shaking of the arm, the length of the operation trajectory of the arm, and the cycle time;

further adding the added reward value for each of the shaking of the arm, the length of the operation trajectory of the arm, and the cycle time together to obtain a further added reward value; and

updating a value function which determines a value of the operation program of the robot based on the state variable, the determination data, and the further added reward value;

determining an operation of the robot based on the learned operation program; and

outputting the determined operation of the robot to the robot,

wherein the machine learning method further comprises:

setting a negative reward among the negative rewards based on the state variable and the determination data and causing the robot to operate with the shaking of the arm being reduced; and

setting a further negative reward among the negative rewards based on the state variable and the determination data and causing the robot to operate with a reduced cycle time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2017
From: TODA, SYUNTAROU
To: FANUC CORPORATION
Reel/Frame 043565/0478 →
Priority Claims (1)
JP JP2016-182233 · Sep 16, 2016 · national
Continuity (1)
Related Publication 20180079076A1 · Mar 22, 2018
Cited By (1)
US 12,427,658