IP Library Granted Patent US 12,377,549
Granted Patent B2
US 12,377,549 · App. 17/833,460 · Granted Aug 5, 2025

System and method for determining a grasping hand model

Inventors: Francesc Moreno Noguer (Barcelona, ES); Guillem Alenyà Ribas (Barcelona, ES); Enric Corona Puyane (Barcelona, ES); Albert Pumarola Peris (Barcelona, ES); Grégory Rogez (Meylan, FR)
Assignee: NAVER Corporation
B25J9/1697B25J9/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,377,549
App. No.
17/833,460
Granted
Aug 5, 2025
Kind
B2
Abstract

Method for determining a grasping hand model suitable for grasping an object by receiving an image including at least one object; obtaining an object model estimating a pose and shape of the object from the image of the object; selecting a grasp class from a set of grasp classes by means of a neural network, with a cross entropy loss, thus, obtaining a set of parameters defining a coarse grasping hand model; refining the coarse grasping hand model, by minimizing loss functions referring to the parameters of the hand model for obtaining an operable grasping hand model while minimizing the distance between the finger of the hand model and the surface of the object and preventing interpenetration; and obtaining a mesh of the hand represented by the enhanced set of parameters.

Claims (282)

1. A system for determining parameters of a hand model suitable for grasping an object, comprising:

a first neural network for segmenting the object in an image;

a second neural network for predicting parameters of the hand model that define a pose for grasping the segmented object; and

a third neural network for refining the predicted parameters of the hand model to fit the segmented object by maximizing a number of contact points between the object in the image and the hand model while minimizing interpenetration;

wherein the refined predicted parameters of the hand model represent how a hand may grasp the object.

2. The system according to claim 1 , wherein said first neural network segments the object in an image by estimating a pose and a 3D shape of the object from the image of the object.

3. The system according to claim 1 , wherein said second neural network predicts parameters of the hand model by predicting a grasp class from a set of grasp classes to obtain a set of parameters defining a coarse grasping hand model.

4. The system according to claim 1 , wherein said third neural network refines the predicted parameters of the hand model by minimizing loss functions referring to the parameters of the hand model for obtaining an operable grasping hand model while minimizing a distance between the fingers of the hand model and a surface of the object and preventing interpenetration.

5. The system according to claim 1 , wherein said second neural network is a Convolutional Neural Network, with a cross entropy loss L class defined as:

L class =Σ c∈K C o,c log(1− P o,c )

wherein C represents a grasp type for the particular object (o), c represents the grasp classes among the K possible grasps classes, and P represents pose predictions for the particular object (o).

6. The system according to claim 1 , wherein said third neural network obtains a representation of a hand grasping the object by using the refined predicted parameters of the hand model, the representation being a mesh of the refined hand model.

7. The system according to claim 1 , wherein said third neural network uses a MANO model, being a 51 degrees of freedom model of a possible human hand, for refining the predicted parameters of the hand model.

8. The system according to claim 1 , wherein said third neural network evaluates the grasping hand model by calculating at least one evaluating metric of an analytical grasp metric, which computes an approximation of the minimum force to be applied to break the grasp stability; an average number of contact fingers, wherein numerous contact points between hand and object favor a strong grasp; a hand-object interpenetration volume, wherein object and hand are voxelized, and the volume shared by both 3D models is computed; a simulation displacement of the object mesh subjected to gravity; and a percentage of graspable objects for which an operable grasp could be predicted, being an operable grasp the one with at least two contact points and no interpenetration.

9. The system according to claim 8 , wherein said third neural network (a) randomly rotates an object model; (b) obtains a grasping hand model for each rotated object model; (c) evaluates each rotated grasping hand model using evaluating metrics; and (d) selects the rotated grasping hand models having the highest score.

10. The system according to claim 1 , wherein said second neural network estimates a pose and shape of the object by using an object reconstruction phase for obtaining a cloud of points representing the object form the obtained image.

11. The system according to claim 1 , wherein said image includes more than one object;

said first neural network segmenting each object in said image and estimating a 3D shape of each segmented object;

said second neural network predicting parameters of the hand model that define a pose for grasping each segmented object;

said third neural network refining the predicted parameters of the hand model to fit each segmented object.

12. The system according to claim 1 , wherein said second neural network selects a grasp class by utilizing a phase of predicting an increment of translation and rotation of the hand model and a modified coarse configuration of the hand model.

13. The system according to claim 1 , wherein said third neural network refines the coarse grasping model by:

(a) selecting at least one articulation (i) of the hand model;

(b) calculating an arc (Ai) between a finger (j) of the hand model and close object vertices (O),

D θ ←min i (min k (∥ A i θ ,O k ∥ 2 ))

(c) estimating the angle the finger needs to be rotated to collide with the object, rotating the articulation for minimizing the arc, thus, reducing the distance between the hand model and the object vertices, including a hyperparameter for controlling the interpenetration of the hand model into the object,

γ j ′←arg min θ D θ +δ, ∀θs.t. D θ <t d

(d) defining the following loss functions:

L

a

r

c

=

1

"\[LeftBracketingBar]"

J

"\[RightBracketingBar]"

j

J

D

θ

j

L

γ

j

J

"\[LeftBracketingBar]"

"\[LeftBracketingBar]"

γ

j

-

γ

j

"\[RightBracketingBar]"

"\[RightBracketingBar]"

2

;

and

(e) minimizing the loss functions defined.

14. The system according to claim 13 , wherein said third neural network refines the coarse grasping model for each articulation sequentially from the knuckle to the tip for each finger.

15. The system according to claim 1 , wherein said third neural network refines the coarse grasping model by minimizing a loss function from a distance between the hand vertices and the target object, wherein is considered that there is a contact when the distance is below to 2 mm, defined by

L

cont

=

1

"\[LeftBracketingBar]"

V

cont

"\[RightBracketingBar]"

v

V

cont

min

k

"\[LeftBracketingBar]"

"\[LeftBracketingBar]"

v

,

O

k

t

"\[RightBracketingBar]"

"\[RightBracketingBar]"

2

.

16. The system according to claim 1 , wherein said third neural network refines the coarse grasping model by minimizing a loss function from a distance of interpenetration between a vertex of the hand model and the object, defined by

L

int

=

1

"\[LeftBracketingBar]"

V

i

"\[RightBracketingBar]"

j

"\[LeftBracketingBar]"

O

"\[RightBracketingBar]"

v

V

i

min

k

v

,

O

k

j

2

.

17. The system according to claim 1 , wherein said third neural network refines the coarse grasping model by minimizing a loss function from a distance below a table plane, between a vertex of the hand model and the table plane, wherein the distance is favored to be positive, defined by L p =Σ v v min (0, |(v−p p )·v p |).

18. The system according to claim 1 , wherein said third neural network refines the coarse grasping model by minimizing a loss function from an adversarial loss function, using a Wasserstein loss including a gradient penalty loss, defined by L adv =−E H,R,T˜p(H,R,T) [D(G(I))]+E H,R,T˜p(H,R,T) [D(H*, R*, T*)].

19. The system according to claim 1 , wherein the hand is a human hand.

20. A method for determining parameters of a hand model suitable for grasping an object, comprising:

(a) receiving an object segmented in an image;

(b) predicting, with a first neural network, parameters of the hand model that define a pose for grasping the segmented object; and

(c) refining, with a second neural network, the predicted parameters of the hand model to fit the segmented object by maximizing a number of contact points between the object in the image and the hand model while minimizing interpenetration;

wherein the refined predicted parameters of the hand model represent how a hand may grasp the object.

21. The method according to claim 20 , wherein said receiving is performed using a third neural network.

22. The method according to claim 20 , wherein (b) predicts parameters of the hand model by predicting a grasp class from a set of grasp classes to obtain a set of parameters defining a coarse grasping hand model.

23. The method according to claim 20 , wherein (c) refines the predicted parameters of the hand model by minimizing loss functions referring to the parameters of the hand model for obtaining an operable grasping hand model while minimizing the distance between the fingers of the hand model and the surface of the object and preventing interpenetration.

24. The method according to claim 20 , further comprising:

(d) obtaining a representation of a hand grasping the object by using the refined hand model.

25. The method according to claim 20 , wherein the first neural network is a Convolutional Neural Network, with a cross entropy loss L class defined as:

L class =Σ c∈K C o,c log (1 −P o,c );

wherein C represents a grasp type for the particular object (o), c represents the grasp classes among the K possible grasps classes, and P represents pose predictions for the particular object (o).

26. The method according to claim 24 , wherein the representation obtained in (d) is a mesh of the refined hand model.

27. The method according to claim 20 , wherein the hand model is represented by using a MANO model, being a 51 degrees of freedom model of a possible human hand.

28. The method according to claim 20 , further comprising:

(d) evaluating the grasping hand model obtained by calculating at least one evaluating metric of an analytical grasp metric, which computes an approximation of the minimum force to be applied to break the grasp stability; an average number of contact fingers, wherein numerous contact points between hand and object favor a strong grasp; a hand-object interpenetration volume, wherein object and hand are voxelized, and the volume shared by both 3D models is computed; a simulation displacement of the object mesh subjected to gravity; and a percentage of graspable objects for which an operable grasp could be predicted, being an operable grasp the one with at least two contact points and no interpenetration.

29. The method according to claim 28 , further comprising:

(e) randomly rotating an object model;

(f) obtaining a grasping hand model for each rotated object model;

(g) evaluating each rotated grasping hand model using evaluating metrics; and

(h) selecting the rotated grasping hand models having the highest score.

30. The method according to claim 21 , wherein estimating a pose and shape of the object comprises an object reconstruction phase for obtaining a cloud of points representing the object form the obtained image.

31. The method according to claim 20 , wherein the image comprises more than one object and the method further comprises the step of repeating (a) to (c) for each object in the image, wherein the objects are known.

32. The method according to claim 20 , wherein (b) selects a grasp class by utilizing a phase of predicting an increment of translation and rotation of the hand model and a modified coarse configuration of the hand model.

33. The method according to claim 20 , wherein (c) includes:

(c1) selecting at least one articulation (i) of the hand model;

(c2) calculating an arc (Ai) between a finger (j) of the hand model and close object vertices (O),

D θ ←min i (min k (∥i A i θ , O k ∥ 2 ));

(c3) estimating the angle the finger needs to be rotated to collide with the object, rotating the articulation for minimizing the arc, thus, reducing the distance between the hand model and the object vertices, including a hyperparameter for controlling the interpenetration of the hand model into the object,

γ j ′←arg min θ D θ +δ,∀θs.t. D θ <t d ;

(c 4 ) defining the following loss functions:

L

arc

=

1

"\[LeftBracketingBar]"

J

"\[RightBracketingBar]"

j

J

D

θ

j

L

γ

j

J

γ

j

-

γ

j

2

;

and

(c5) minimizing the loss functions defined.

34. The method according to claim 33 , wherein (c) further includes (c6) repeating phases (c2) and (c3) for each articulation sequentially from the knuckle to the tip for each finger.

35. The method according to claim 20 , wherein said (c) further comprises minimizing a loss function from a distance between the hand vertices and the target object, wherein is considered that there is a contact when the distance is below to 2 mm, defined by

L

cont

=

1

"\[LeftBracketingBar]"

V

cont

"\[RightBracketingBar]"

v

V

cont

min

k

v

,

O

k

t

2

.

36. The method according to claim 20 , wherein said (c) further comprises minimizing a loss function from a distance of interpenetration between a vertex of the hand model and the object, defined by

L

int

=

1

"\[LeftBracketingBar]"

V

i

"\[RightBracketingBar]"

j

"\[LeftBracketingBar]"

O

"\[RightBracketingBar]"

v

V

i

min

k

v

,

O

k

j

2

.

37. The method according to claim 20 , wherein said (c) further comprises minimizing a loss function from a distance below a table plane, between a vertex of the hand model and the table plane, wherein the distance is favored to be positive, defined by L p =Σ v v min (0, |(v−p p )·v p |).

38. The method according to claim 20 , wherein said (c) further comprises minimizing a loss function from an adversarial loss function, using a Wasserstein loss including a gradient penalty loss, defined by L adv =−E H,R,T˜p(H,R,T) [D(G(I))]+E H,R,T˜p(H,R,T) [D(H*, R*, T*)].

39. The method according to claim 20 , wherein the hand is a human hand.

40. The method according to claim 20 , wherein (a) segments an object in an image by estimating a pose and shape of the object from the image of the object.

Assignments (5)
CONFIRMATORY ASSIGNMENT Recorded Sep 29, 2025
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 072979/0392 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2024
From: NAVERS LAB CORPORATION
To: NAVER CORPORATION
Reel/Frame 068789/0417 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: MORENO NOGUER, FRANCESC; ALENYÀ RIBAS, GUILLEM; CORONA PUYANE, ENRIC; PUMAROLA PERIS, ALBERT; ROGEZ, GRÉGORY
To: NAVER FRANCE
Reel/Frame 060493/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: NAVER FRANCE
To: NAVER CORPORATION
Reel/Frame 060493/0419 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2022
From: NAVER CORPORATION
To: NAVER LABS CORPORATION
Reel/Frame 060493/0474 →
Priority Claims (1)
ES 202030553 · Jun 9, 2020 · national
Continuity (3)
Continuation In Part 17341970 · Jun 8, 2021
Provisional Application 63208231 · Jun 8, 2021
Related Publication 20220402125A1 · Dec 22, 2022
References Cited (106)
US 10099369B2 · Kopicki · 2018 [cited by examiner]
US 11341406B1 · Redmon · 2022 [cited by examiner]
US 20140016856A1 · Jiang · 2014 [cited by examiner]
US 20200086483A1 · Li · 2020 [cited by examiner]
US 20210122045A1 · Handa · 2021 [cited by examiner]
Li, Xueting, Sifei Liu, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, and Jan Kautz. “Putting Humans in a Scene: Learning Affordance in 3D Indoor Environments.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern R… [cited by applicant]
Lin, Yun, and Yu Sun. “Grasp Planning Based on Strategy Extracted from Demonstration.” In 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems, 4458-63. Chicago, IL, USA: IEEE, 2014. [cited by applicant]
Liu, Jia, Fangxiaoyu Feng, Yuzuko C. Nakamura, and Nancy S. Pollard. “A Taxonomy of Everyday Grasps in Action.” In 2014 IEEE-RAS International Conference on Humanoid Robots, 573-80. Madrid, Spain: IEEE, 2014. [cited by applicant]
Lopes, Manuel, Francisco Melo, Luis Montesano, and José Santos-Victor. “Abstraction Levels for Robotic Imitation: Overview and Computational Approaches.” In From Motor Learning to Interaction Learning in Robots, edited … [cited by applicant]
Madadi, Meysam, Sergio Escalera, Xavier Baro, and Jordi Gonzalez. “End-to-End Global to Local CNN Learning for Hand Pose Recovery in Depth Data.” arXiv, Apr. 11, 2018. [cited by applicant]
Mahler, Jeffrey, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg. “Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Gras… [cited by applicant]
Miller, Andew and Allen, Peter—Graspit! A versatile simulator for robotic grasping. 2004. [cited by applicant]
Minjie Cai, Kris M. Kitani, and Yoichi Sato. “A Scalable Approach for Understanding the Visual Structures of Hand Grasps.” In 2015 IEEE International Conference on Robotics and Automation (ICRA), 1360-66. Seattle, WA, U… [cited by applicant]
Morrison, Douglas, Peter Corke, and Jürgen Leitner. “Closing the Loop for Robotic Grasping: A Real-Time, Generative Grasp Synthesis Approach.” arXiv, May 15, 2018. [cited by applicant]
Mousavian, Arsalan, Clemens Eppner, and Dieter Fox. “6-DOF GraspNet: Variational Grasp Generation for Object Manipulation.” In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2901-10. Seoul, Korea (Sou… [cited by applicant]
Mueller, Franziska, Florian Bernard, Oleksandr Sotnychenko, Dushyant Mehta, Srinath Sridhar, Dan Casas, and Christian Theobalt. “GANerated Hands for Real-Time 3D Hand Tracking from Monocular RGB.” In 2018 IEEE/CVF Confe… [cited by applicant]
Myronenko, Andriy, and Xubo Song. “Point-Set Registration: Coherent Point Drift.” arXiv, May 15, 2009. [cited by applicant]
Oberweger, Markus, and Vincent Lepetit. “DeepPrior++: Improving Fast and Accurate 3D Hand Pose Estimation.” In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), 585-94. Venice: IEEE, 2017. [cited by applicant]
Oberweger, Markus, Paul Wohlhart, and Vincent Lepetit. “Generalized Feedback Loop for Joint Hand-Object Pose Estimation.” arXiv, Mar. 25, 2019. [cited by applicant]
Oikonomidis, lason, Nikolaos Kyriazis, and Antonis A. Argyros. “Full DOF Tracking of a Hand Interacting with an Object by Modeling Occlusions and Physical Constraints.” In 2011 International Conference on Computer Visio… [cited by applicant]
Panteleris, Paschalis, and Antonis Argyros. “Back to RGB: 3D Tracking of Hands and Hand-Object Interactions Based on Short-Baseline Stereo.” In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), 57… [cited by applicant]
Panteleris, Paschalis, Iason Oikonomidis, and Antonis Argyros. “Using a Single RGB Frame for Real Time 3D Hand Pose Estimation in the Wild.” In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), 436-… [cited by applicant]
Pas, Andreas ten, Marcus Gualtieri, Kate Saenko, and Robert Platt. “Grasp Pose Detection in Point Clouds.” arXiv, Jun. 29, 20. [cited by applicant]
Pham, Tu-Hoa, Nikolaos Kyriazis, Antonis A. Argyros, and Abderrahmane Kheddar. “Hand-Object Contact Force Estimation from Markerless Visual Tracking.” IEEE Transactions on Pattern Analysis and Machine Intelligence 40, N… [cited by applicant]
Pinto, Lerrel, and Abhinav Gupta. “Supersizing Self-Supervision: Learning to Grasp from 50K Tries and 700 Robot Hours.” In 2016 IEEE International Conference on Robotics and Automation (ICRA), 3406-13. Stockholm, Sweden… [cited by applicant]
Porzi, Lorenzo et al.—Learning Depth-Aware Deep representations for robotic perception. IEEE Robotics and Automation Letters, 2016. [cited by applicant]
Pumarola, Albert, Jordi Sanchez, Gary P. T. Choi, Alberto Sanfeliu, and Francesc Moreno. “3DPeople: Modeling the Geometry of Dressed Humans.” In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2242-51.… [cited by applicant]
Qian, Chen, Xiao Sun, Yichen Wei, Xiaoou Tang, and Jian Sun. “Realtime and Robust Hand Tracking from Depth.” In 2014 IEEE Conference on Computer Vision and Pattern Recognition, 1106-13. Columbus, OH, USA: IEEE, 2014. [cited by applicant]
Rad, Mahdi, Markus Oberweger, and Vincent Lepetit. “Domain Transfer for 3D Pose Estimation from Color Images without Manual Annotations.” arXiv, Feb. 21, 2019. [cited by applicant]
Redmon, Joseph, and Anelia Angelova. “Real-Time Grasp Detection Using Convolutional Neural Networks.” In 2015 IEEE International Conference on Robotics and Automation (ICRA), 1316-22. Seattle, WA, USA: IEEE, 2015. [cited by applicant]
Rock, Jason, Tanmay Gupta, Justin Thorsen, JunYoung Gwak, Daeyun Shin, and Derek Hoiem. “Completing 3D Object Shape from One Depth Image.” In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2484-… [cited by applicant]
Rogez, Gregory, James S. Supancic, and Deva Ramanan. “First-Person Pose Recognition Using Egocentric Workspaces.” In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4325-33. Boston, MA, USA: IEEE… [cited by applicant]
Rogez, Gregory, James S. Supancic, and Deva Ramanan. “Understanding Everyday Hands in Action from RGB-D Images.” In 2015 IEEE International Conference on Computer Vision (ICCV), 3889-97. Santiago: IEEE, 2015. [cited by applicant]
Romero, Javier, Hedvig Kjellstrom, and Danica Kragic. “Hands in Action: Real-Time 3D Reconstruction of Hands in Interaction with Objects.” In 2010 IEEE International Conference on Robotics and Automation, 458-63. Anchor… [cited by applicant]
Romero, Javier, Dimitrios Tzionas, and Michael J. Black. “Embodied Hands: Modeling and Capturing Hands and Bodies Together.” ACM Transactions on Graphics 36, No. 6—Nov. 20, 2017. [cited by applicant]
Roy, Anirban, and Sinisa Todorovic. “A Multi-Scale CNN for Affordance Segmentation in RGB Images.” In Computer Vision—ECCV 2016, edited by Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, 9908:186-201. Lecture Not… [cited by applicant]
Saran, Akanksha, Damien Teney, and Kris M. Kitani. “Hand Parsing for Fine-Grained Recognition of Human Grasps in Monocular Images.” In 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 505… [cited by applicant]
Saudabayev, Artur, Zhanibek Rysbek, Raykhan Khassenova, and Huseyin Atakan Varol. “Human Grasping Database for Activities of Daily Living with Depth, Color and Kinematic Data Streams.” Scientific Data 5, No. 1. 2018. [cited by applicant]
Schmidt, Philipp, Nikolaus Vahrenkamp, Mirko Wachter, and Tamim Asfour. “Grasping of Unknown Objects Using Deep Convolutional Neural Networks Based on Depth Images.” In 2018 IEEE International Conference on Robotics and… [cited by applicant]
Spurr, Adrian, Jie Song, Seonwook Park, and Otmar Hilliges. “Cross-Modal Deep Variational Hand Pose Estimation.” In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 89-98. Salt Lake City, UT: IEEE, 2… [cited by applicant]
Sridhar, Srinath, Franziska Mueller, Michael Zollhofer, Dan Casas, Antti Oulasvirta, and Christian Theobalt. “Real-Time Joint Tracking of a Hand Manipulating an Object from RGB-D Input.” In Computer Vision—ECCV 2016, ed… [cited by applicant]
Sun, Xiao, Yichen Wei, Shuang Liang, Xiaoou Tang, and Jian Sun. “Cascaded Hand Pose Regression,” CVPR, 2015. [cited by applicant]
Sundermeyer, Martin, Zoltan-Csaba Marton, Maximilian Durner, Manuel Brucker, and Rudolph Triebel. “Implicit 3D Orientation Learning for 6D Object Detection from RGB Images.” In Computer Vision—ECCV 2018, edited by Vitto… [cited by applicant]
Supancic, James S., Gregory Rogez, Yi Yang, Jamie Shotton, and Deva Ramanan. “Depth-Based Hand Pose Estimation: Data, Methods, and Challenges.” In 2015 IEEE International Conference on Computer Vision (ICCV), 1868-76. S… [cited by applicant]
Tekin, Bugra, Federica Bogo, and Marc Pollefeys. “H+O: Unified Egocentric Recognition of 3D Hand-Object Poses and Interactions.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4506-15. Lo… [cited by applicant]
Thayananthan, A., B. Stenger, P.H.S. Torr, and R. Cipolla. “Shape Context and Chamfer Matching in Cluttered Scenes.” In 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings… [cited by applicant]
Tsoli, Aggeliki, and Antonis A. Argyros. “Joint 3D Tracking of a Deformable Object in Interaction with a Hand.” In Computer Vision—ECCV 2018, edited by Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair W… [cited by applicant]
Tzionas, Dimitrios, Luca Ballan, Abhilash Srikantha, Pablo Aponte, Marc Pollefeys, and Juergen Gall. “Capturing Hands in Action Using Discriminative Salient Points and Physics Simulation.” International Journal of Compu… [cited by applicant]
Tzionas, Dimitrios, and Juergen Gall. “3D Object Reconstruction from Hand-Object Interactions.” In 2015 IEEE International Conference on Computer Vision (ICCV), 729-37. Santiago, Chile: IEEE, 2015. [cited by applicant]
Wang, Xiaolong, Rohit Girdhar, and Abhinav Gupta. “Binge Watching: Scaling Affordance Learning from Sitcoms.” In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3366-75. Honolulu, HI: IEEE, 2017. [cited by applicant]
Wang, Yangang, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu, Qionghai Dai, and Jinxiang Chai. “Video-Based Hand Manipulation Capture through Composite Motion Control.” ACM Transactions on Graphics 32, No. 4—Jul. 21, … [cited by applicant]
Wang, Yiwei, Xin Ye, Yezhou Yang, and Wenlong Zhang. “Hand Movement Prediction Based Collision-Free Human-Robot Interaction.” In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 492-93.… [cited by applicant]
Xiang, Yu, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. “PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes.” In Robotics: Science and Systems XIV. Robotics: Science and … [cited by applicant]
Yuan, Shanxin, Guillermo Garcia-Hernando, Bjorn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov, et al. “Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals.” In 2018 IE… [cited by applicant]
Yuan, Shanxin, Qi Ye, Bjorn Stenger, Siddhant Jain, and Tae-Kyun Kim. “BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis.” In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), … [cited by applicant]
Zimmermann, Christian, and Thomas Brox. “Learning to Estimate 3D Hand Pose from Single RGB Images.” In 2017 International Conference on Computer Vision (ICCV), 4913-21. Venice: IEEE, 2017. [cited by applicant]
Arjovsky, Martin, Soumith Chintala, and Léon Bottou. “Wasserstein GAN.” arXiv, Dec. 6, 2017. [cited by applicant]
Baek, Seungryul, Kwang In Kim, and Tae-Kyun Kim. “Pushing the Envelope for RGB-Based Dense 3D Hand Pose Estimation via Neural Rendering.” arXiv, Apr. 9, 2019. [cited by applicant]
Balasubramanian, Ravi, Ling Xu, Peter D. Brook, Joshua R. Smith, and Yoky Matsuoka. “Physical Human Interactive Guidance: Identifying Grasping Principles From Human-Planned Grasps.” IEEE Transactions on Robotics 28, No.… [cited by applicant]
Ballan, Luca, Aparna Taneja, Jürgen Gall, Luc Van Gool, and Marc Pollefeys. “Motion Capture of Hands in Action Using Discriminative Salient Points.” In Computer Vision—ECCV 2012, edited by Andrew Fitzgibbon, Svetlana La… [cited by applicant]
Bambach, Sven, Stefan Lee, David J. Crandall, and Chen Yu. “Lending A Hand: Detecting Hands and Recognizing Activities in Complex Egocentric Interactions.” In 2015 IEEE International Conference on Computer Vision (ICCV)… [cited by applicant]
Borras, Júlia, Guillem Alenya, and Carme Torras. “A Grasping-Centered Analysis for Cloth Manipulation.” arXiv, Apr. 9, 2020. [cited by applicant]
Boukhayma, Adnane, Rodrigo de Bem, and Philip H.S. Torr. “3D Hand Shape and Pose From Images in the Wild.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10835-44. Long Beach, CA, USA: IE… [cited by applicant]
Brahmbhatt, Samarth, Cusuh Ham, Charles C. Kemp, and James Hays. “ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging.” arXiv, Apr. 14, 2019. [cited by applicant]
Bullock, Ian M., Thomas Feix, and Aaron M. Dollar. “The Yale Human Grasping Dataset: Grasp, Object, and Task Data in Household and Machine Shop Environments.” The International Journal of Robotics Research 34, No. 3 (Ma… [cited by applicant]
Cai, Minjie, Kris Kitani, and Yoichi Sato. “Understanding Hand-Object Manipulation by Modeling the Contextual Relationship between Actions, Grasp Types and Object Attributes.” arXiv, Jul. 22, 2018. [cited by applicant]
Cai, Yujun, Liuhao Ge, Jianfei Cai, and Junsong Yuan. “Weakly-Supervised 3D Hand Pose Estimation from Monocular RGB Images.” In Computer Vision—ECCV 2018, edited by Vittorio Ferrari, Martial Hebert, Cristian Sminchisesc… [cited by applicant]
Calli, Berk, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M. Dollar. “The YCB Object and Model Set: Towards Common Benchmarks for Manipulation Research.” In 2015 International Conference on… [cited by applicant]
Chang, Angel X., Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, et al. “ShapeNet: An Information-Rich 3D Model Repository.” arXiv, Dec. 9, 2015. [cited by applicant]
Chang, Ju Yong, Gyeongsik Moon, and Kyoung Mu Lee. “V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map.” In 2018 IEEE/CVF Conference on Computer Vision … [cited by applicant]
Choi, Chiho, Sang Ho Yoon, Chin-Ning Chen, and Karthik Ramani. “Robust Hand Pose Estimation during the Interaction with an Unknown Object.” In 2017 IEEE International Conference on Computer Vision (ICCV), 3142-51. Venic… [cited by applicant]
Corona, Enric, Albert Pumarola, Guillem Alenya, and Francesc Moreno-Noguer. “Context-Aware Human Motion Prediction.” arXiv, Mar. 23, 2020. [cited by applicant]
Corona, Enric et al.—Pose Estimation for Objects with rotational symmetry. IROS. 2018. [cited by applicant]
Cutkosky, M.R. “On Grasp Choice, Grasp Models, and the Design of Hands for Manufacturing Tasks.” IEEE Transactions on Robotics and Automation 5, No. 3 (Jun. 1989): 269-79. [cited by applicant]
De-An Huang, Minghuang Ma, Wei-Chiu Ma, and Kris M. Kitani. “How Do We Use Our Hands? Discovering a Diverse Set of Common Grasps.” In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 666-75. Bosto… [cited by applicant]
Do, Thanh-Toan, Anh Nguyen, and lan Reid. “AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection.” In 2018 IEEE International Conference on Robotics and Automation (ICRA), 1-5. Brisbane, QL… [cited by applicant]
Edwards, S.J et al—Developmental and functional hand grasps. Slcak 2002. [cited by applicant]
Ferrari, C. and Canny, J. F. Planning Optimal Grasps. ICRA, 1992. [cited by applicant]
Feix, Thomas, Roland Pawlik, Heinz-Bodo Schmiedmayer, Javier Romero, and Danica Kragic. “A Comprehensive Grasp Taxonomy,” RSS, 2009. [cited by applicant]
Feix, Thomas, Javier Romero, Heinz-Bodo Schmiedmayer, Aaron M. Dollar, and Danica Kragic. “The GRASP Taxonomy of Human Grasp Types.” IEEE Transactions on Human-Machine Systems 46, No. 1 (Feb. 2016). [cited by applicant]
Garcia-Hernando, Guillermo, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. “First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations.” In 2018 IEEE/CVF Conference on Computer Vision and Pattern… [cited by applicant]
Ge, Liuhao, Hui Liang, Junsong Yuan, and Daniel Thalmann. “Robust 3D Hand Pose Estimation in Single Depth Images: From Single-View CNN to Multi-View CNNs.” In 2016 IEEE Conference on Computer Vision and Pattern Recognit… [cited by applicant]
Ge, Liuhao, Zhou Ren, Yuncheng Li, Zehao Xue, Yingying Wang, Jianfei Cai, and Junsong Yuan. “3D Hand Shape and Pose Estimation From a Single RGB Image,” CVPR, 2019. [cited by applicant]
Groueix, Thibault, Matthew Fisher, Vladimir Kim, Bryan Russell, and Mathieu Aubry. “AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation,” CVPR, 2018. [cited by applicant]
Gualtieri, Marcus, Andreas ten Pas, Kate Saenko, and Robert Platt. “High Precision Grasp Pose Detection in Dense Clutter.” arXiv, Jun. 22, 2017. [cited by applicant]
Gulrajani, Ishaan, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. “Improved Training of Wasserstein GANs,” NIPS, 2017. [cited by applicant]
Gupta, Abhinav, Scott Satkin, Alexei A. Efros, and Martial Hebert. “From 3D Scene Geometry to Human Workspace.” In CVPR 2011, 1961-68. Colorado Springs, CO, USA: IEEE, 2011. [cited by applicant]
Hamer, Henning, Juergen Gall, Thibaut Weise, and Luc Van Gool. “An Object-Dependent Hand Pose Prior from Sparse Training Data.” In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 671-78… [cited by applicant]
Hamer, Henning, Konrad Schindler, Esther Koller-Meier, and Luc Van Gool. “Tracking a Hand Manipulating an Object.” In 2009 IEEE 12th International Conference on Computer Vision, 1475-82. Kyoto, Japan: IEEE, 2009. [cited by applicant]
Hampali, Shreyas, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. “HOnnotate: A Method for 3D Annotation of Hand and Object Poses.” arXiv, May 30, 2020. [cited by applicant]
Hasson, Yana, Gul Varol, Dimitrios Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. “Learning Joint Reconstruction of Hands and Manipulated Objects.” In 2019 IEEE/CVF Conference on Computer … [cited by applicant]
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition.” In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770-78. Las Vegas, NV, USA: IEEE, 2016. [cited by applicant]
Hu, Yinlin, Joachim Hugonot, Pascal Fua, and Mathieu Salzmann. “Segmentation-Driven 6D Object Pose Estimation.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3380-89. Long Beach, CA, USA… [cited by applicant]
Iqbal, Umar, Pavlo Molchanov, Thomas Breuel, Juergen Gall, and Jan Kautz. “Hand Pose Estimation via Latent 2.5D Heatmap Regression.” In Computer Vision—ECCV 2018, edited by Vittorio Ferrari, Martial Hebert, Cristian Smi… [cited by applicant]
Keskin, Cem, Furkan Kraç, Yunus Emre Kara, and Lale Akarun. “Hand Pose Estimation and Hand Shape Classification Using Multi-Layered Randomized Decision Forests.” In Computer Vision—ECCV 2012, edited by Andrew Fitzgibbon… [cited by applicant]
Kim, Min Ku, Ramviyas Nattanmai Parasuraman, Liu Wang, Yeonsoo Park, Bongjoong Kim, Seung Jun Lee, Nanshu Lu, Byung-Cheol Min, and Chi Hwan Lee. “Soft-Packaged Sensory Glove System for Human-like Natural Interaction and… [cited by applicant]
Kokic, Mia, Danica Kragic, and Jeannette Bohg. “Learning Task-Oriented Grasping from Human Activity Datasets.” arXiv, 2019. [cited by applicant]
Kokic, Mia, Danica Kragic, and Jeannette Bohg. “Learning to Estimate Pose and Shape of Hand-Held Objects from RGB Images.” arXiv, Nov. 11, 2019. [cited by applicant]
Kumra, Sulabh, and Christopher Kanan. “Robotic Grasp Detection Using Deep Convolutional Neural Networks.” arXiv, Jul. 21, 2017. [cited by applicant]
La Gorce, M. de, D. J. Fleet, and N. Paragios. “Model-Based 3D Hand Pose Estimation from Monocular Video.” IEEE Transactions on Pattern Analysis and Machine Intelligence 33, No. 9 (Sep. 2011. [cited by applicant]
Lai, Kevin, Liefeng Bo, Xiaofeng Ren, and Dieter Fox. “A Large-Scale Hierarchical Multi-View RGB-D Object Dataset,” ICRA, 2011. [cited by applicant]
Lau, Manfred, Kapil Dev, Weiqi Shi, Julie Dorsey, and Holly Rushmeier. “Tactile Mesh Saliency.” ACM Transactions on Graphics. 2016. [cited by applicant]
Lei, Jie, Mingli Song, Ze-Nian Li, and Chun Chen. “Whole-Body Humanoid Robot Imitation with Pose Similarity Evaluation.” Signal Processing. 2015. [cited by applicant]
Lenz, Ian, Honglak Lee, and Ashutosh Saxena. “Deep Learning for Detecting Robotic Grasps.” The International Journal of Robotics Research 34, No. 4-5—Apr. 2015. [cited by applicant]
Lepetit, Vincent, Francesc Moreno-Noguer, and Pascal Fua. “EPnP: An Accurate O(n) Solution to the PnP Problem.” International Journal of Computer Vision 81, No. 2 (Feb. 2009. [cited by applicant]
Levine, Sergey, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen. “Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection.” The International Journal of Ro… [cited by applicant]
Cited By (1)
US 12,657,764