IP Library › Granted Patent US 10,635,944
Granted Patent B2
US 10,635,944 · App. 16/443,765 · Granted Apr 28, 2020

Self-supervised robotic object interaction

Inventors: Eric Victor Jang (Cupertino, CA); Sergey Vladimir Levine (Berkeley, CA); Coline Manon Devin (Berkeley, CA)
Assignee: Google LLC
G06K9/6262G06K9/00664G06K9/6256G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,635,944
App. No.
16/443,765
Granted
Apr 28, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training an object representation neural network. One of the methods includes obtaining training sets of images, each training set comprising: (i) a before image of a before scene of the environment, (ii) an after image of an after scene of the environment after the robot has removed a particular object, and (iii) an object image of the particular object, and training the object representation neural network on the batch of training data, comprising determining an update to the object representation parameters that encourages the vector embedding of the particular object in each training set to be closer to a difference between (i) the vector embedding of the after scene in the training set and (ii) the vector embedding of the before scene in the training set.

Claims (90)

1. A method of training an object interaction task neural network that (i) has a plurality of object interaction parameters and (ii) is used to select actions to be performed by a robot to cause the robot to perform a task that includes performing a specified interaction with a particular object of interest in an environment conditioned on an image of the particular object of interest, the method comprising:

obtaining a goal object image of a goal object selected from a plurality of objects currently located in the environment;

processing the goal object image using an object representation neural network having a plurality of object representation parameters, wherein the object representation neural network is configured to process the goal object image in accordance with current values of the plurality of object representation parameters to generate a vector embedding of the goal object;

controlling the robot to perform an episode of the task by selecting actions to be performed by the robot using the object interaction task neural network while the object interaction task neural network is conditioned on the vector embedding of the goal object and in accordance with current values of the plurality of object interaction parameters;

generating, from the actions performed during the episode, a sequence of state representation—action pairs, the state representation in each state representation—action pair characterizing a state of the environment when the action in the state representation—action pair was performed by the robot during the episode;

determining whether the robot successfully performed the task for any of the plurality of objects in the environment during the episode;

when the robot successfully performed the task for any of the plurality of objects in the environment:

determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment, comprising:

obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed;

processing the successful object image using the object representation neural network in accordance with the current values of the object representation parameters to generate a vector embedding of the successful object;

determining a similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

determining a first reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

for each reward value of the one or more reward values, training the object interaction task neural network using the sequence of state representation—action pairs and the reward value of the one or more reward values.

2. The method of claim 1 , further comprising, when the robot did not successfully perform the task for any of the plurality of objects:

determining a fourth reward value that indicates that the robot failed at performing the task; and

training the object interaction task neural network using the sequence of state representation—action pairs and the fourth reward value.

3. The method of claim 1 , wherein determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment comprises:

setting a second reward value to a value that indicates that the task was successfully completed, and

wherein training the object interaction task neural network using the sequence of state representation—action pairs and the second reward value comprises assigning the one of the plurality of objects for which the task was successfully performed as the goal object for the training of the object interaction task neural network.

4. The method of claim 3 , further comprising:

selecting an alternate object in the environment that is different from the goal object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and a third reward value that indicates that the robot failed at performing the task and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

5. The method of claim 1 , wherein obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed comprises:

causing the robot to place the one of the plurality of objects in a field of view of a camera; and

capturing an image of the one of the plurality of objects using the camera.

6. The method of claim 1 further comprising:

selecting an alternate object in the environment that is different from the goal object;

determining a similarity measure between the vector embedding of the successful object and a vector embedding of the alternate object;

determining a fifth reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the alternate object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and the fifth reward value and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

7. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for training an object interaction task neural network that (i) has a plurality of object interaction parameters and (ii) is used to select actions to be performed by a robot to cause the robot to perform a task that includes performing a specified interaction with a particular object of interest in an environment conditioned on an image of the particular object of interest, the operations comprising:

obtaining a goal object image of a goal object selected from a plurality of objects currently located in the environment;

processing the goal object image using an object representation neural network having a plurality of object representation parameters, wherein the object representation neural network is configured to process the goal object image in accordance with current values of the plurality of object representation parameters to generate a vector embedding of the goal object;

controlling the robot to perform an episode of the task by selecting actions to be performed by the robot using the object interaction task neural network while the object interaction task neural network is conditioned on the vector embedding of the goal object and in accordance with current values of the plurality of object interaction parameters;

generating, from the actions performed during the episode, a sequence of state representation—action pairs, the state representation in each state representation—action pair characterizing a state of the environment when the action in the state representation—action pair was performed by the robot during the episode;

determining whether the robot successfully performed the task for any of the plurality of objects in the environment during the episode;

when the robot successfully performed the task for any of the plurality of objects in the environment:

determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment, comprising:

obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed;

processing the successful object image using the object representation neural network in accordance with the current values of the object representation parameters to generate a vector embedding of the successful object;

determining a similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

determining a first reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

for each reward value of the one or more reward values, training the object interaction task neural network using the sequence of state representation—action pairs and the reward value of the one or more reward values.

8. The system of claim 7 , the operations further comprising, when the robot did not successfully perform the task for any of the plurality of objects:

determining a fourth reward value that indicates that the robot failed at performing the task; and

training the object interaction task neural network using the sequence of state representation—action pairs and the fourth reward value.

9. The system of claim 7 , wherein determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment comprises:

setting a second reward value to a value that indicates that the task was successfully completed, and

wherein training the object interaction task neural network using the sequence of state representation—action pairs and the second reward value comprises assigning the one of the plurality of objects for which the task was successfully performed as the goal object for the training of the object interaction task neural network.

10. The system of claim 7 , the operations further comprising:

selecting an alternate object in the environment that is different from the goal object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and a third reward value that indicates that the robot failed at performing the task and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

11. The system of claim 7 , wherein obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed comprises:

causing the robot to place the one of the plurality of objects in a field of view of a camera; and

capturing an image of the one of the plurality of objects using the camera.

12. The system of claim 7 the operations further comprising:

selecting an alternate object in the environment that is different from the goal object;

determining a similarity measure between the vector embedding of the successful object and a vector embedding of the alternate object;

determining a fifth reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the alternate object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and the fifth reward value and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

13. One or more non-transitory computer-readable storage media storing instructions that are operable, when executed by one or more computers, to cause the one or more computers to perform operations for training an object interaction task neural network that (i) has a plurality of object interaction parameters and (ii) is used to select actions to be performed by a robot to cause the robot to perform a task that includes performing a specified interaction with a particular object of interest in an environment conditioned on an image of the particular object of interest, the operations comprising:

obtaining a goal object image of a goal object selected from a plurality of objects currently located in the environment;

processing the goal object image using an object representation neural network having a plurality of object representation parameters, wherein the object representation neural network is configured to process the goal object image in accordance with current values of the plurality of object representation parameters to generate a vector embedding of the goal object;

controlling the robot to perform an episode of the task by selecting actions to be performed by the robot using the object interaction task neural network while the object interaction task neural network is conditioned on the vector embedding of the goal object and in accordance with current values of the plurality of object interaction parameters;

generating, from the actions performed during the episode, a sequence of state representation—action pairs, the state representation in each state representation—action pair characterizing a state of the environment when the action in the state representation—action pair was performed by the robot during the episode;

determining whether the robot successfully performed the task for any of the plurality of objects in the environment during the episode;

when the robot successfully performed the task for any of the plurality of objects in the environment:

determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment, comprising:

obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed;

processing the successful object image using the object representation neural network in accordance with the current values of the object representation parameters to generate a vector embedding of the successful object;

determining a similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

determining a first reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the goal object; and

for each reward value of the one or more reward values, training the object interaction task neural network using the sequence of state representation—action pairs and the reward value of the one or more reward values.

14. The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising, when the robot did not successfully perform the task for any of the plurality of objects:

determining a fourth reward value that indicates that the robot failed at performing the task; and

training the object interaction task neural network using the sequence of state representation—action pairs and the fourth reward value.

15. The one or more non-transitory computer-readable storage media of claim 13 , wherein determining one or more reward values based on the robot successfully performing the task for one of the plurality of objects in the environment comprises:

setting a second reward value to a value that indicates that the task was successfully completed, and

wherein training the object interaction task neural network using the sequence of state representation—action pairs and the second reward value comprises assigning the one of the plurality of objects for which the task was successfully performed as the goal object for the training of the object interaction task neural network.

16. The one or more non-transitory computer-readable storage media of claim 13 , the operations further comprising:

selecting an alternate object in the environment that is different from the goal object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and a third reward value that indicates that the robot failed at performing the task and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

17. The one or more non-transitory computer-readable storage media of claim 13 , wherein obtaining a successful object image of the one of the plurality of objects for which the task was successfully performed comprises:

causing the robot to place the one of the plurality of objects in a field of view of a camera; and

capturing an image of the one of the plurality of objects using the camera.

18. The one or more non-transitory computer-readable storage media of claim 13 the operations further comprising:

selecting an alternate object in the environment that is different from the goal object;

determining a similarity measure between the vector embedding of the successful object and a vector embedding of the alternate object;

determining a fifth reward value based on the similarity measure between the vector embedding of the successful object and the vector embedding of the alternate object; and

training the object interaction task neural network (i) using the sequence of state representation—action pairs and the fifth reward value and (ii) with the alternate object assigned as the goal object for the training of the object interaction task neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2019
From: JANG, ERIC VICTOR; LEVINE, SERGEY VLADIMIR; DEVIN, COLINE MANON
To: GOOGLE LLC
Reel/Frame 049609/0927 →
Continuity (2)
Provisional Application 62685885 · Jun 15, 2018
Related Publication 20190385022A1 · Dec 19, 2019
Cited By (1)
US 12,493,792