IP Library Granted Patent US 11,559,885
Granted Patent B2
US 11,559,885 · App. 17/375,798 · Granted Jan 24, 2023

Method and system for grasping an object

Inventors: Ahmad Humayun (Union City, CA); Michael Stark (Union City, CA); Nan Rong (Union City, CA); Bhaskara Mannar Marthi (Union City, CA); Aravind Sivakumar (Union City, CA)
Assignee: Intrinsic Innovation LLC
B25J9/1612B25J9/161B25J9/1664B25J19/023G06N3/08G06T7/50G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,559,885
App. No.
17/375,798
Granted
Jan 24, 2023
Kind
B2
Abstract

The method for increasing the accuracy of grasping an object can include: labelling an image based on an attempted object grasp by a robot and generating a trained graspability network using the labelled images. The method can additionally or alternatively include determining a grasp point using the trained graspability network; executing an object grasp at the grasp point S 400 ; and/or any other suitable elements.

Claims (70)

1. A method comprising:

generating training data for training a neural network configured to, after processing an input image, generate a probability map indicating respective likelihoods of successfully grasping a corresponding object at a grasp point by a robotic manipulator, wherein generating the training data comprises:

for each of a plurality of object scenes:

receiving an image capturing the object scene of the plurality of object scenes, wherein the object scene comprises a first object and a second object that is at least partially occluded by the first object in the imagee;

determining object parameters for the first object and the second object within the object scene using an object detector, wherein the object parameters represent information of a corresponding object in the image;

determining a set of candidate grasp points for the image based on the object parameters, each of the set of candidate grasp points associated with a corresponding object of the first object and the second object;

selecting, according to one or more criteria, a grasp point from the set of candidate grasp points for grasping a corresponding object;

receiving grasp outcome data representing a grasp outcome for the grasp point, the grasp outcome data being based on whether the corresponding object was successfully grasped by the robotic manipulator; and

updating the image including labeling the grasp point in the image with the grasp outcome; and

training the neural network, using the generated training data that includes the updated images capturing the plurality of object scenes.

2. The method of claim 1 , wherein the information for a corresponding object comprises at least one or more of: a keypoint associated with the corresponding object, a bounding box for the corresponding object, a label for the corresponding object, a pose of the corresponding object, an object mask for the corresponding object, or a visibility score for the corresponding object.

3. The method of claim 2 , wherein the neural network is a fully convolutional neural network (FCN).

4. The method of claim 1 , wherein the one or more criteria comprise:

a proximity associated with a candidate grasp point to an edge of a container for the object scene, a level of occlusion between the first object and the second object, a depth value associated with the candidate grasp point, a keypoint type associated with the grasp point, or a keypoint label associated with the grasp point.

5. The method of claim 1 , wherein the first object and the second object are homogeneous.

6. The method of claim 1 , further comprising:

receiving an additional image capturing an object scene different from the plurality of object scenes;

processing the additional image using the neural network, to generate a probability map for the additional image,

selecting an exploratory grasp point for the additional image based on the probability map and an exploration rule;

determining a grasp outcome for the exploratory grasp point; and

updating the neural network using the additional image, the exploratory grasp point, and the grasp outcome for the exploratory grasp point.

7. The method of claim 1 , wherein the object detector is trained using a set of artificially-generated training examples capturing multiple object scenes.

8. The method of claim 1 , wherein generating the training data further comprises:

for the image capturing the object scene, determining a grasp pose for the robotic manipulator based on the object parameters, wherein the grasp pose for the robotic manipulator represents a pose for the robotic manipulator to grasp the corresponding object at the grasp point, and

labeling the grasp point in the image with the grasp pose.

9. The method of claim 1 , wherein the probability map for the input image comprises a respective probability of grasp success for each of at least a majority of pixels of the input image.

10. The method of claim 1 , wherein the neural network comprises a pretrained depth enhancement network, trained based on images depicting different scenes, noisy depth information associated with each different scene, and accurate depth information associated with each different scene; wherein the pretrained depth enhancement network outputs accurate depth information for a test image, wherein the images depicting different scenes are different from the images capturing the plurality of object scenes.

11. The method of claim 1 , wherein determining a grasp outcome for the grasp point comprises controlling a robotic manipulator to grasp the corresponding object at the grasp point.

12. The method of claim 11 , wherein the training data further comprises a manipulator pose of the robotic manipulator for each grasp point in the updated images; wherein the probability map for the input image further comprises a predicted manipulator pose for a corresponding robotic manipulator grasping at a predicted grasp point associated with an object depicted within the input image.

13. The method of claim 11 , wherein the training data further comprises a manipulator pose of the robotic manipulator for each grasp point in the updated images; wherein the probability map for the input image further comprises a predicted manipulator pose a corresponding robotic manipulator grasping at each pixel of input image.

14. A system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform respective operations, the operations comprising:

generating training data for training a neural network configured to, after processing an input image, generate a probability map indicating respective likelihoods of successfully grasping a corresponding object at a grasp point by a robotic manipulator, wherein generating the training data comprises:

for each of a plurality of object scenes:

receiving an image capturing the object scene of the plurality of object scenes, wherein the object scene comprises a first object and a second object that is at least partially occluded by the first object in the image;

determining object parameters for the first object and the second object within the object scene using an object detector, wherein the object parameters represent information of a corresponding object in the image;

determining a set of candidate grasp points for the image based on the object parameters, each of the set of candidate grasp points associated with a corresponding object of the first object and the second object;

selecting, according to one or more criteria, a grasp point from the set of candidate grasp points for grasping a corresponding object;

receiving grasp outcome data representing a grasp outcome for the grasp point, the grasp outcome data being based on whether the corresponding object was successfully grasped by the robotic manipulator; and

updating the image including labeling the grasp point in the image with the grasp outcome; and

training the neural network, using the generated training data that includes the updated images capturing the plurality of object scenes.

15. The system of claim 14 , wherein the information for a corresponding object comprises at least one or more of: a keypoint associated with the corresponding object, a bounding box for the corresponding object, a label for the corresponding object, a pose of the corresponding object, an object mask for the corresponding object, or a visibility score for the corresponding obj ect.

16. The system of claim 14 , wherein the operations further comprise:

receiving an additional image capturing an object scene different from the plurality of object scenes;

processing the additional image using the neural network, to generate a probability map for the additional image,

selecting an exploratory grasp point for the additional image based on the probability map and an exploration rule;

determining a grasp outcome for the exploratory grasp point; and

updating the neural network using the additional image, the exploratory grasp point, and the grasp outcome for the exploratory grasp point.

17. The system of claim 14 , wherein generating the training data further comprises:

for the image capturing the object scene, determining a grasp pose for the robotic manipulator based on the object parameters, wherein the grasp pose for the robotic manipulator represents a pose for the robotic manipulator to grasp the corresponding object at the grasp point, and

labeling the grasp point in the image with the grasp pose.

18. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform respective operations, the respective operations comprising:

generating training data for training a neural network configured to, after processing an input image, generate a probability map indicating respective likelihoods of successfully grasping a corresponding object at a grasp point by a robotic manipulator, wherein generating the training data comprises:

for each of a plurality of object scenes:

receiving an image capturing the object scene of the plurality of object scenes, wherein the object scene comprises a first object and a second object that is at least partially occluded by the first object in the image;

determining object parameters for the first object and the second object within the object scene using an object detector, wherein the object parameters represent information of a corresponding object in the image;

determining a set of candidate grasp points for the image based on the object parameters, each of the set of candidate grasp points associated with a corresponding object of the first object and the second object;

selecting, according to one or more criteria, a grasp point from the set of candidate grasp points for grasping a corresponding object;

receiving grasp outcome data representing a grasp outcome for the grasp point, the grasp outcome data being based on whether the corresponding object was successfully grasped by the robotic manipulator; and

updating the image including labeling the grasp point in the image with the grasp outcome; and

training the neural network, using the generated training data that includes the updated images capturing the plurality of object scenes.

19. The one or more non-transitory computer-readable storage media of claim 18 , wherein the information for a corresponding object comprises at least one or more of: a keypoint associated with the corresponding object, a bounding box for the corresponding object, a label for the corresponding object, a pose of the corresponding object, an object mask for the corresponding object, or a visibility score for the corresponding object.

20. The one or more non-transitory computer-readable storage media of claim 18 , wherein the operations further comprise:

receiving an additional image capturing an object scene different from the plurality of object scenes;

processing the additional image using the neural network, to generate a probability map for the additional image,

selecting an exploratory grasp point for the additional image based on the probability map and an exploration rule;

determining a grasp outcome for the exploratory grasp point; and

updating the neural network using the additional image, the exploratory grasp point, and the grasp outcome for the exploratory grasp point.

21. The one or more non-transitory computer-readable storage media of claim 18 , wherein generating the training data further comprises:

for the image capturing the object scene, determining a grasp pose for the robotic manipulator based on the object parameters, wherein the grasp pose for the robotic manipulator represents a pose for the robotic manipulator to grasp the corresponding object at the grasp point, and

labeling the grasp point in the image with the grasp pose.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE RECEIVING PARTY NAME PREVIOUSLY RECORDED AT REEL: 060389 FRAME: 0682. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 7, 2022
From: VICARIOUS FPC, INC.; BOSTON POLARIMETRICS, INC.
To: INTRINSIC INNOVATION LLC
Reel/Frame 060614/0104 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: VICARIOUS FPC, INC; BOSTON POLARIMETRICS, INC.
To: LLC, INTRINSIC I
Reel/Frame 060389/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2021
From: HUMAYUN, AHMAD; STARK, MICHAEL; RONG, NAN; MARTHI, BHASKARA MANNAR; SIVAKUMAR, ARAVIND
To: VICARIOUS FPC, INC.
Reel/Frame 056967/0105 →
Continuity (4)
Provisional Application 63164078 · Mar 22, 2021
Provisional Application 63162360 · Mar 17, 2021
Provisional Application 63051844 · Jul 14, 2020
Related Publication 20220016766A1 · Jan 20, 2022
Cited By (3)
US 12,466,078 US 12,521,888 US 12,691,586