IP Library › Granted Patent US 12,521,888
Granted Patent B2
US 12,521,888 · App. 18/367,827 · Granted Jan 13, 2026

Synergies between pick and place: task-aware grasp estimation

Inventors: Nikhil Narsingh Chavan Dafle (Jersey City, NJ); Vasileios Vasilopoulos (Woodbridge, NJ); Shubham Agrawal (Jersey City, NJ); Jinwook Huh (Millburn, NJ); Suveer Garg (New York, NY); Pedro Piacenza (Jersey City, NJ); Isaac Hisano Kasahara (Brooklyn, NY); Kazim Selim Engin (Weehawken, NJ); Zhanpeng He (New York, NY); Shuran Song (New York, NY); Ibrahim Volkan Isler (Saint Paul, MN)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
B25J9/1697B25J9/1612B25J9/163G06T7/60G06T7/73G06T2207/10024G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,521,888
App. No.
18/367,827
Granted
Jan 13, 2026
Kind
B2
Abstract

Systems, methods, and apparatuses for controlling a robot including a manipulator, including: determining three-dimensional (3D) geometry information about a target object based on an image of the target object; determining 3D geometry information about a scene in which the target object is to be placed based on at least one image of the scene; obtaining affordance information by providing the 3D geometry information about the target object and the 3D geometry information about the scene to at least one neural network model; commanding the robot to grasp the target object using the manipulator according to a grasp orientation corresponding to the affordance information; and commanding the robot to position the manipulator according to a placement direction corresponding to the affordance information in order to place the target object at a location in the scene.

Claims (60)

1 . An electronic device for controlling a robot including a manipulator, the electronic device comprising:

one or more processors configured to:

determine three-dimensional (3D) geometry information about a target object based on an image of the target object;

determine 3D geometry information about a scene in which the target object is to be placed based on at least one image of the scene;

determine a plurality of candidate placement directions for placing the target object in the scene based on the 3D geometry information about the scene;

obtain affordance information by providing the 3D geometry information about the target object and information about the plurality of candidate placement directions to at least one neural network model;

command the robot to grasp the target object using the manipulator according to a grasp orientation corresponding to the affordance information; and

command the robot to position the manipulator according to a placement direction corresponding to the affordance information in order to place the target object at a location in the scene.

2 . The electronic device of claim 1 , wherein the one or more processors are further configured to:

determine a plurality of candidate grasp orientations for grasping the target object based on the 3D geometry information about the target object;

obtain a plurality of affordance maps by providing information about the plurality of candidate grasp orientations and the information about the plurality of candidate placement directions to the at least one neural network model, wherein each affordance map from among the plurality of affordance maps corresponds to a candidate grasp orientation from among the plurality of candidate grasp orientations and a candidate placement direction from among the plurality of candidate placement directions; and

select an affordance map from among the plurality of affordance maps, wherein the affordance information corresponds to the selected affordance map.

3 . The electronic device of claim 2 , wherein the at least one neural network model comprises:

an object encoder configured to output a plurality of object encodings corresponding to the plurality of candidate grasp orientations based on the 3D geometry information about the target object;

a scene encoder configured to output a plurality of scene encodings corresponding to the plurality of candidate placement directions based on the 3D geometry information about the scene; and

an affordance decoder configured to output the plurality of affordance maps based on the plurality of object encodings and the plurality of scene encodings.

4 . The electronic device of claim 3 , wherein the object encoder, the scene encoder, and the affordance decoder are jointly trained.

5 . The electronic device of claim 2 , wherein the affordance map comprises a plurality of pixels corresponding to a plurality of affordance values, and

wherein each affordance value from among the plurality of affordance values indicates a probability of success for placing the target object.

6 . The electronic device of claim 5 , wherein the affordance map is selected based on the plurality of pixels including a highest affordance value from among all affordance values associated with the plurality of affordance maps.

7 . The electronic device of claim 1 , further comprising at least one camera configured to capture the image of the target object and the at least one image of the scene.

8 . The electronic device of claim 7 , wherein the image of the target object is a depth image, and

wherein the at least one image of the scene is a color image.

9 . The electronic device of claim 1 , wherein the one or more processors are configured to command the robot to position the manipulator by computing a proposed trajectory based on the placement direction, and generating a velocity command corresponding to the proposed trajectory.

10 . A method for controlling a robot including a manipulator, the method comprising:

determining three-dimensional (3D) geometry information about a target object based on an image of the target object;

determining 3D geometry information about a scene in which the target object is to be placed based on at least one image of the scene;

determining a plurality of candidate placement directions for placing the target object in the scene based on the 3D geometry information about the scene;

obtaining affordance information by providing the 3D geometry information about the target object and information about the plurality of candidate placement directions to at least one neural network model;

commanding the robot to grasp the target object using the manipulator according to a grasp orientation corresponding to the affordance information; and

commanding the robot to position the manipulator according to a placement direction corresponding to the affordance information in order to place the target object at a location in the scene.

11 . The method of claim 10 , further comprising:

determining a plurality of candidate grasp orientations for grasping the target object based on the 3D geometry information about the target object;

obtaining a plurality of affordance maps by providing information about the plurality of candidate grasp orientations and the information about the plurality of candidate placement directions to the at least one neural network model, wherein each affordance map from among the plurality of affordance maps corresponds to a candidate grasp orientation from among the plurality of candidate grasp orientations and a candidate placement direction from among the plurality of candidate placement directions; and

selecting an affordance map from among the plurality of affordance maps,

wherein the affordance information corresponds to the selected affordance map.

12 . The method of claim 11 , wherein the at least one neural network model comprises:

an object encoder configured to output a plurality of object encodings corresponding to the plurality of candidate grasp orientations based on the 3D geometry information about the target object;

a scene encoder configured to output a plurality of scene encodings corresponding to the plurality of candidate placement directions based on the 3D geometry information about the scene; and

an affordance decoder configured to output the plurality of affordance maps based on the plurality of object encodings and the plurality of scene encodings.

13 . The method of claim 12 , wherein the object encoder, the scene encoder, and the affordance decoder are jointly trained.

14 . The method of claim 11 , wherein the affordance map comprises a plurality of pixels corresponding to a plurality of affordance values, and

wherein each affordance value from among the plurality of affordance values indicates a probability of success for placing the target object.

15 . The method of claim 14 , wherein the affordance map is selected based on the plurality of pixels including a highest affordance value from among all affordance values associated with the plurality of affordance maps.

16 . The method of claim 10 , further comprising capturing the image of the target object and the at least one image of the scene.

17 . The method of claim 16 , wherein the image of the target object is a depth image, and

wherein the at least one image of the scene is a color image.

18 . The method of claim 10 , wherein the commanding the robot to position the manipulator comprises computing a proposed trajectory based on the placement direction, and generating a velocity command corresponding to the proposed trajectory.

19 . A non-transitory computer-readable medium configured to store instructions which, when executed by at least one processor of a device for controlling a robot including a manipulator, cause the at least one processor to:

determine three-dimensional (3D) geometry information about a target object based on an image of the target object;

determine 3D geometry information about a scene in which the target object is to be placed based on at least one image of the scene;

determine a plurality of candidate placement directions for placing the target object in the scene based on the 3D geometry information about the scene;

obtain affordance information by providing the 3D geometry information about the target object and information about the plurality of candidate placement directions to at least one neural network model;

command the robot to grasp the target object using the manipulator according to a grasp orientation corresponding to the affordance information; and

command the robot to position the manipulator according to a placement direction corresponding to the affordance information in order to place the target object at a location in the scene.

20 . The non-transitory computer-readable medium of claim 19 , the instructions further cause the at least one processor to:

determine a plurality of candidate grasp orientations for grasping the target object based on the 3D geometry information about the target object; and

obtain a plurality of affordance maps by providing information about the plurality of candidate grasp orientations and the information about the plurality of candidate placement directions to the at least one neural network model, wherein each affordance map from among the plurality of affordance maps corresponds to a candidate grasp orientation from among the plurality of candidate grasp orientations and a candidate placement direction from among the plurality of candidate placement directions; and

select an affordance map from among the plurality of affordance maps,

wherein the affordance information corresponds to the selected affordance map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: CHAVAN DAFLE, NIKHIL NARSINGH; VASILOPOULOS, VASILEIOS; AGRAWAL, SHUBHAM; HUH, JINWOOK; GARG, SUVEER; PIACENZA, PEDRO; KASAHARA, ISAAC HISANO; ENGIN, KAZIM SELIM; HE, ZHANPENG; SONG, SHURAN; ISLER, IBRAHIM VOLKAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064893/0860 →
Continuity (4)
Provisional Application 63452620 · Mar 16, 2023
Provisional Application 63450908 · Mar 8, 2023
Provisional Application 63406853 · Sep 15, 2022
Related Publication 20240091951A1 · Mar 21, 2024
References Cited (23)
US 7859540B2 · Dariush · 2010 [cited by applicant]
US 9616568B1 · Russell · 2017 [cited by applicant]
US 9987746B2 · Bradski et al. · 2018 [cited by applicant]
US 10131051B1 · Goyal · 2018 [cited by examiner]
US 11559885B2 · Humayun et al. · 2023 [cited by applicant]
US 11597078B2 · Yang et al. · 2023 [cited by applicant]
US 20190022863A1 · Kundu et al. · 2019 [cited by applicant]
US 20220168899A1 · Boroushaki et al. · 2022 [cited by applicant]
US 20220391638A1 · Fan · 2022 [cited by applicant]
US 20220392162A1 · Shen et al. · 2022 [cited by applicant]
US 20240335941A1 · Aparicio Ojea · 2024 [cited by examiner]
CA 3203956A1 · 2022 [cited by applicant]
EP 2952301B1 · 2019 [cited by applicant]
JP 202161014A · 2021 [cited by applicant]
KR 102067878B1 · 2020 [cited by applicant]
WO 2021198053A1 · 2021 [cited by applicant]
WO 2021206671A1 · 2021 [cited by applicant]
WO 2022119962A1 · 2022 [cited by applicant]
International Search Report (PCT/ISA/210) issued on Dec. 13, 2023 by the International Searching Authority for International Patent Application No. PCT/KR2023/013966. [cited by applicant]
Communication issued Jun. 26, 2025 by the European Patent Office in European patent Application No. 23865911.4. [cited by applicant]
Manuelli, Lucas et al., “kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation”, arXiv:1903.06684v2 [cs.RO], Oct. 29, 2019. (26 pages total). [cited by applicant]
Deng, Yuhong et al., “Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered Environment”, 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, Nov. 4-8, 2019,… [cited by applicant]
Danielczuk, Michael et al., “Object Rearrangement Using Learned Implicit Collision Functions”, arXiv:2011.107261 [cs.RO], Nov. 21, 2020. (8 pages total). [cited by applicant]