IP Library Granted Patent US 12,243,238
Granted Patent B1
US 12,243,238 · App. 18/224,551 · Granted Mar 4, 2025

Hand pose estimation for machine learning based gesture recognition

Inventors: Jonathan Marsden (San Mateo, CA); Raffi Bedikian (San Francisco, CA); David Samuel Holz (San Francisco, CA)
Assignee: ULTRAHAPTICS IP TWO LIMITED
G06T7/13G06T2207/10028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,238
App. No.
18/224,551
Granted
Mar 4, 2025
Kind
B1
Abstract

The technology disclosed performs hand pose estimation on a so-called “joint-by-joint” basis. So, when a plurality of estimates for the 28 hand joints are received from a plurality of expert networks (and from master experts in some high-confidence scenarios), the estimates are analyzed at a joint level and a final location for each joint is calculated based on the plurality of estimates for a particular joint. This is a novel solution discovered by the technology disclosed because nothing in the field of art determines hand pose estimates at such granularity and precision. Regarding granularity and precision, because hand pose estimates are computed on a joint-by-joint basis, this allows the technology disclosed to detect in real time even the minutest and most subtle hand movements, such a bend/yaw/tilt/roll of a segment of a finger or a tilt an occluded finger, as demonstrated supra in the Experimental Results section of this application.

Claims (68)

1. A method of preparing sample hand positions for training of neural network systems, the method including:

accessing simulation parameters that specify at least one of:

a range of hand positions and position sequences,

a range of hand anatomies, including palm size, fattiness, stubbiness, and skin tone, and

a range of backgrounds;

accessing a camera perspective specification that specifies one or more of:

a focal length,

a field of view of the camera,

a wavelength sensitivity, and

artificial lighting conditions;

generating a plurality of hand position-hand anatomy-background simulations, each simulation labeled with hand position parameters, including joint locations of joints of the hand in three dimensions (3D), in a ground truth vector for training a convolutional neural network, the simulations organized in sequences;

applying the camera perspective specification to render from the simulations at least a corresponding set of simulated hand position images; and

saving the simulated hand position images with one or more labelled hand position parameters from the corresponding simulations.

2. The method of claim 1 , wherein the simulated hand position images are stereoscopic images with depth map information.

3. The method of claim 1 , wherein the simulated hand position images are binocular pairs of images.

4. The method of claim 1 , wherein the hand position parameters are a plurality of joint locations in three-dimensional (3D) space.

5. The method of claim 1 , wherein the hand position parameters are a plurality of joint angles in three-dimensional (3D) space.

6. The method of claim 1 , wherein the hand position parameters are a plurality of hand skeleton segments in three-dimensional (3D) space.

7. The method of claim 1 , the method further including:

extracting stereoscopic hand boundaries for the simulated hand position images and aligning the hand boundaries with hand centers included in the simulated hand position images;

generating translated, rotated and scaled variants of the hand boundaries and applying Gaussian jittering to the variants;

extracting hand regions from one or more jittered variants of the hand boundaries;

computing ground truth pose vectors for the hand regions; and

storing the pose vectors in tangible machine readable memory as output labels for the simulated hand position images.

8. The method of claim 7 , further including generating three-dimensional (3D) simulated hands in mesh models and/or capsule hand skeleton models.

9. The method of claim 7 , further including generating a simulated coordinate system to determine hand position parameters of a simulated hand in three-dimensional (3D).

10. The method of claim 7 , further including generating a simulated perspective of a simulated gesture recognition system to determine hand position parameters of a simulated hand in three-dimensional (3D).

11. The method of claim 7 , wherein the ground truth pose vectors are 84 dimensional representing 28 hand joints in three-dimensional (3D) space.

12. A non-transitory computer readable storage medium impressed with computer program instructions to prepare sample hand positions for training of neural network systems, the instructions, when executed on a processor, implement a method comprising:

accessing simulation parameters that specify at least one of:

a range of hand positions and position sequences,

a range of hand anatomies, including palm size, fattiness, stubbiness, and skin tone, and

a range of backgrounds;

accessing a camera perspective specification that specifies one or more of:

a focal length,

a field of view of the camera,

a wavelength sensitivity, and

artificial lighting conditions;

generating a plurality of hand position-hand anatomy-background simulations, each simulation labeled with hand position parameters, including joint locations of joints of the hand in three dimensions (3D), in a ground truth vector for training a convolutional neural network, the simulations organized in sequences;

applying the camera perspective specification to render from the simulations at least a corresponding set of simulated hand position images; and

saving the simulated hand position images with one or more labelled hand position parameters from the corresponding simulations.

13. The non-transitory computer readable storage medium of claim 12 , wherein the simulated hand position images are stereoscopic images with depth map information.

14. The non-transitory computer readable storage medium of claim 12 , wherein the simulated hand position images are binocular pairs of images.

15. The non-transitory computer readable storage medium of claim 12 , wherein the hand position parameters include parameters selected from one or more of: (i) a plurality of joint locations in three-dimensional (3D) space; (ii) a plurality of joint angles in three-dimensional (3D) space; and (iii) a plurality of hand skeleton segments in three-dimensional (3D) space.

16. The non-transitory computer readable storage medium of claim 12 , further including instructions which when executed by the processor, implement:

extracting stereoscopic hand boundaries for the simulated hand position images and aligning the hand boundaries with hand centers included in the simulated hand position images;

generating translated, rotated and scaled variants of the hand boundaries and applying Gaussian jittering to the variants;

extracting hand regions from one or more jittered variants of the hand boundaries;

computing ground truth pose vectors for the hand regions; and

storing the pose vectors in tangible machine readable memory as output labels for the simulated hand position images.

17. The non-transitory computer readable storage medium of claim 16 , further including generating three-dimensional (3D) simulated hands in mesh models and/or capsule hand skeleton models.

18. The non-transitory computer readable storage medium of claim 16 , further including generating a simulated coordinate system to determine hand position parameters of a simulated hand in three-dimensional (3D).

19. The non-transitory computer readable storage medium of claim 16 , further including generating a simulated perspective of a simulated gesture recognition system to determine hand position parameters of a simulated hand in three-dimensional (3D).

20. The non-transitory computer readable storage medium of claim 16 , wherein the ground truth pose vectors are 84 dimensional representing 28 hand joints in three-dimensional (3D) space.

21. A system including one or more processors coupled to memory, the memory loaded with computer instructions, which instructions, when executed on the one or more processors, implement actions comprising the method of claim 1 :

accessing simulation parameters that specify at least one of:

a range of hand positions and position sequences,

a range of hand anatomies, including palm size, fattiness, stubbiness, and skin tone, and

a range of backgrounds;

accessing a camera perspective specification that specifies one or more of:

a focal length,

a field of view of the camera,

a wavelength sensitivity, and

artificial lighting conditions;

generating a plurality of hand position-hand anatomy-background simulations, each simulation labeled with hand position parameters, including joint locations of joints of the hand in three dimensions (3D), in a ground truth vector for training a convolutional neural network, the simulations organized in sequences;

applying the camera perspective specification to render from the simulations at least a corresponding set of simulated hand position images; and

saving the simulated hand position images with one or more labelled hand position parameters from the corresponding simulations for use in training a hand position recognition system.

22. The method of claim 1 , wherein the plurality of hand position-hand anatomy-background simulations includes simulations representing at least two different hands with different respective hand anatomies.

Assignments (9)
SECURITY INTEREST Recorded Apr 6, 2026
From: SIM IP HXR LLC
To: UNITY MASTER LLC SERIES XIX
Reel/Frame 075365/0907 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2026
From: ULTRAHAPTICS IP TWO LIMITED
To: SIM IP HXR LLC
Reel/Frame 075127/0604 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2026
From: ULTRAHAPTICS LIMITED; ULTRAHAPTICS IP LIMITED; ULTRAHAPTICS IP TWO LIMITED; ULTRALEAP LIMITED
To: SIM IP HXR LLC
Reel/Frame 074403/0864 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: LEAP MOTION, INC.
To: LMI LIQUIDATING CO. LLC
Reel/Frame 069885/0367 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: MARSDEN, JONATHAN
To: OCUSPEC
Reel/Frame 069883/0904 →
CHANGE OF NAME Recorded Jan 15, 2025
From: OCUSPEC, INC.
To: LEAP MOTION, INC.
Reel/Frame 069932/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: LMI LIQUIDATING CO. LLC
To: ULTRAHAPTICS IP TWO LIMITED
Reel/Frame 069885/0526 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: BEDIKIAN, RAFFI
To: LEAP MOTION, INC.
Reel/Frame 069884/0188 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2025
From: HOLZ, DAVID
To: OCUSPEC, INC.
Reel/Frame 069885/0032 →
Continuity (4)
Division 16508231 · Jul 10, 2019
Continuation 15432872 · Feb 14, 2017
Provisional Application 62335534 · May 12, 2016
Provisional Application 62296561 · Feb 17, 2016
References Cited (125)
US 4990838A · Kawato et al. · 1991 [cited by applicant]
US 5659764A · Sakiyama et al. · 1997 [cited by applicant]
US 8854433B1 · Rafii · 2014 [cited by applicant]
US 8971572B1 · Yin et al. · 2015 [cited by applicant]
US 9002099B2 · Litvak et al. · 2015 [cited by applicant]
US 9058663B2 · Andriluka et al. · 2015 [cited by applicant]
US 9383895B1 · Vinayak et al. · 2016 [cited by applicant]
US 9442564B1 · Dillon · 2016 [cited by applicant]
US 9501716B2 · Fleishman et al. · 2016 [cited by applicant]
US 9690984B2 · Butler et al. · 2017 [cited by applicant]
US 10048765B2 · Tang et al. · 2018 [cited by applicant]
US 10156603B1 · Fu · 2018 [cited by examiner]
US 10733381B2 · Fuchizaki · 2020 [cited by applicant]
US 11841920B1 · Marsden et al. · 2023 [cited by applicant]
US 11854308B1 · Marsden et al. · 2023 [cited by applicant]
US 11874970B2 · Bedikian · 2024 [cited by examiner]
US 20020036617A1 · Pryor · 2002 [cited by applicant]
US 20060087510A1 · Adamo-Villani et al. · 2006 [cited by applicant]
US 20080094351A1 · Nogami et al. · 2008 [cited by applicant]
US 20090102845A1 · Takemoto et al. · 2009 [cited by applicant]
US 20090103780A1 · Nishihara et al. · 2009 [cited by applicant]
US 20100235786A1 · Maizels et al. · 2010 [cited by applicant]
US 20120117514A1 · Kim et al. · 2012 [cited by applicant]
US 20120268364A1 · Minnen · 2012 [cited by applicant]
US 20120309532A1 · Ambrus et al. · 2012 [cited by applicant]
US 20130253705A1 · Goldfarb et al. · 2013 [cited by applicant]
US 20130257734A1 · Marti et al. · 2013 [cited by applicant]
US 20130259391A1 · Kawaguchi et al. · 2013 [cited by applicant]
US 20130329011A1 · Lee et al. · 2013 [cited by applicant]
US 20130336524A1 · Zhang et al. · 2013 [cited by applicant]
US 20140079300A1 · Wolfer et al. · 2014 [cited by applicant]
US 20140201690A1 · Holz · 2014 [cited by examiner]
US 20140232631A1 · Fleischmann et al. · 2014 [cited by applicant]
US 20140267666A1 · Holz · 2014 [cited by applicant]
US 20140363076A1 · Han et al. · 2014 [cited by applicant]
US 20150023607A1 · Babin et al. · 2015 [cited by applicant]
US 20150054850A1 · Tanaka · 2015 [cited by examiner]
US 20150077326A1 · Kramer et al. · 2015 [cited by applicant]
US 20150153833A1 · Pinault et al. · 2015 [cited by applicant]
US 20150193124A1 · Schwesinger et al. · 2015 [cited by applicant]
US 20150253863A1 · Babin et al. · 2015 [cited by applicant]
US 20150278589A1 · Mazurenko et al. · 2015 [cited by applicant]
US 20150290795A1 · Oleynik · 2015 [cited by applicant]
US 20150310629A1 · Utsunomiya et al. · 2015 [cited by applicant]
US 20150331493A1 · Algreatly · 2015 [cited by applicant]
US 20160018985A1 · Bennet et al. · 2016 [cited by applicant]
US 20160125243A1 · Arata et al. · 2016 [cited by applicant]
US 20160171340A1 · Fleishman et al. · 2016 [cited by applicant]
US 20160196672A1 · Chertok et al. · 2016 [cited by applicant]
US 20160241791A1 · Narayanswamy · 2016 [cited by examiner]
US 20160246369A1 · Osman · 2016 [cited by applicant]
US 20160259417A1 · Gu · 2016 [cited by applicant]
US 20160313798A1 · Connor · 2016 [cited by applicant]
US 20170060254A1 · Molchanov et al. · 2017 [cited by applicant]
US 20170168586A1 · Sinha et al. · 2017 [cited by applicant]
US 20170177077A1 · Yang et al. · 2017 [cited by applicant]
US 20170192514A1 · Karmon et al. · 2017 [cited by applicant]
US 20170193288A1 · Freedman et al. · 2017 [cited by applicant]
US 20170193289A1 · Karmon et al. · 2017 [cited by applicant]
US 20170206405A1 · Molchanov et al. · 2017 [cited by applicant]
US 20170278304A1 · Hildreth et al. · 2017 [cited by applicant]
US 20170329403A1 · Lai · 2017 [cited by applicant]
US 20180024641A1 · Mao et al. · 2018 [cited by applicant]
US 20180039334A1 · Cohen et al. · 2018 [cited by applicant]
US 20180067545A1 · Provancher et al. · 2018 [cited by applicant]
US 20180101247A1 · Lee et al. · 2018 [cited by applicant]
US 20180101520A1 · Fuchizaki · 2018 [cited by applicant]
US 20180205621A1 · Ungar · 2018 [cited by examiner]
US 20190147233A1 · Cherveny et al. · 2019 [cited by applicant]
Chua et al. (“Model-based 3D hand posture estimation from a single 2D image”, Image and Vision Computing, 2001). (Year: 2001). [cited by examiner]
Ahmed, “A Neural Network based Real Time Hand Gesture Recognition System”, Dec. 2012, 6 pages. [cited by applicant]
Chan, “PCANet: A Simple Deep Learning Baseline for Image Classification?”, Aug. 28, 2014, 15 pages. [cited by applicant]
Chen, “Automatic Generation of Statistical Pose and Shape Models for Atriculated Joints”, Feb. 2, 2014, 12 pages. [cited by applicant]
Choi, “A Collaborative Filtering Approach to Real-Time Hand Pose Estimation”, 2015, 9 pages. [cited by applicant]
“CS231n Convolutional Neural Networks for Visual Recognition”, Mar. 28, 2016, 16 pages. [cited by applicant]
Gibiansky, “Convolutional Neural Networks”, Feb. 24, 2014, 7 pages. [cited by applicant]
Girshick, “Region-based Convolutional Networks for Accurate Object Detection and Segmentation”, May 25, 2015, 16 pages. [cited by applicant]
Han, “Space-Time Representation of People Based on 3D Skeletal Data”, Jan. 21, 2016, 20 pages. [cited by applicant]
Hasan, “Static hand gesture recognition using neural networks”, Jan. 12, 2012, 36 pages. [cited by applicant]
Hijazi, “Using Convolutional Neural Networks for Image Recognition”, 2015, 12 pages. [cited by applicant]
Huang, “Large-scale Learning with SVM and Convolutional Nets”, 2006, 8 pages. [cited by applicant]
Hussein, “Human Action Recognition Using a Temporal Hierarchy of Covariance Descriptors on 3D Joint Locations”, 2013, 7 pages. [cited by applicant]
Ibraheem, “Vision Based Gesture Recognition Using Neural Networks Approaches: A Review”, 2012, 14 pages. [cited by applicant]
Karlgaard, “Adaptive Huber-Based Filtering Using Projection Statistics”, Aug. 18-21, 2008, 21 pages. [cited by applicant]
Knutsson, “Hand Detection and Pose Estimation using Convolutional Neural Networks”, 2015, 118 pages. [cited by applicant]
Krizhevsky, “ImageNet Classification with Deep Convolutional Neural Networks”, 2012, 9 pages. [cited by applicant]
Liu, “Implementation of Training Convolutional Neural Networks”, Jun. 4, 2015, 10 pages. [cited by applicant]
Li, “Fast and Robust Method for Dynamic Gesture Recognition Using Hermite Neural Network”, May 2012, 6 pages. [cited by applicant]
Li, “Heterogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network”, Jun. 13, 2014, 8 pages. [cited by applicant]
McCartney, “Gesture Recognition with the Leap Motion Controller”, 2015, 7 pages. [cited by applicant]
Molchanov, “Hand Gesture Recognition with 3D Convolutional Neural Networks”, Jun. 2015, 7 pages. [cited by applicant]
Oberweger, “Hands Deep in Deep Learning for Hand Pose Estimation”, Feb. 9-11, 2015, 10 pages. [cited by applicant]
Oikonomidis, Efficient Model-based 3D Tracking of Hand Articulations using Kinect, 2011, 11 pgs. [cited by applicant]
Pfister, “Deep Convolutional Neural Networks for Efficient Pose Estimation in Gesture Videos”, Dec. 2013, 16 pages. [cited by applicant]
Presti, “3D Skeletion-based Human Action Classification”, 2011, 29 pages. [cited by applicant]
Qian, “Realtime and Robust Hand Tracking from Depth”, 8 pages. [cited by applicant]
Sharp, “Accurate, Robust, and Flexible Realtime Hand Tracking”, Apr. 18-23, 2015, 10 pages. [cited by applicant]
Sinha, “DeepHand: Robust Hand Pose Estimation by Completing a Matrix Imputed with Deep Features”, Jun. 2016, 9pgs. [cited by applicant]
Socher, “Convolutional-Recursive Deep Learning for 3D Object Classification”, 2012, 9 pages. [cited by applicant]
Supancic III, “Depth-based hand pose estimation: methods, data, and challenges”, May 6, 2015, 15 pages. [cited by applicant]
Tompson, “Real-Time Continuous Pose Recovery of Human Hands Using Convolutional Networks”, Aug. 2014, 11 pages. [cited by applicant]
“Using Neural Nets to Recognize Handwritten Digits”, Mar. 28, 2016, 55 pages. [cited by applicant]
Wang, “Human Action Recognition with Depth Cameras” Chapter 2, 2014, 31 pages. [cited by applicant]
Xu, “Efficient Hand Pose Estimation from a Single Depth Image”, 2013, 7 pages. [cited by applicant]
Zheng, “A Project on Gesture Recognition with Neural Networks for ‘Introduction to Artificial Intelligence’ Classes”, 2010, 14 pages. [cited by applicant]
Sharp, Toby, et al., “Accurate, Robust, and Flexible Real-Time Hand Tracking”, Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, Apr. 18, 2015, pp. 3633-3642. [cited by applicant]
Mekala, et al., “Real-time Sign Language Recognition based on Neural Network Architecture”, 2011 IEEE 43rd Southeastern Symposium on System Theory, Mar. 14-16, 2011, pp. 195-199. [cited by applicant]
Tang et al. , “Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose”, 2015 IEEE International Conference on Computer Vision p. 3325-3333 (Year: 2015). [cited by applicant]
Athitsos et al. “Estimating 3D Hand Pose from a Cluttered Image”, 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'03) (Year: 2003). [cited by applicant]
Xu et al. “Estimate Hand Poses Efficiently from Single Depth Images”, Int J Comput Vis (2016) 116:21-45 (Year: 2016). [cited by applicant]
Chua et al. “Model based 3D hand posure estimation from a single 2D image”, 2002 Elsevier Science (Year: 2002). [cited by applicant]
Erol et al. “Vision-based hand pose estimation: A review”, Computer Vision and Image Understanding 108 (2007) 52-73 (Year: 2007). [cited by applicant]
Khamis et al. “Learning an Efficient Model of Hand Shape Variation from Depth Images”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (Year: 2015). [cited by applicant]
Krejov et al. “Combining Discriminative and Model Based Approaches for Hand Pose Estimation”, 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG) (Year: 2015). [cited by applicant]
Oberweger et al. “Training a Feedback Loop for Hand Pose Estimation”, 2015 IEEE International Conference on Computer Vision (ICCV) (Year: 2015). [cited by applicant]
Oberweger et al. “Deep Prior++: Improving Fast and Accurate 3D Hand Pose Estimation”, ICCV Workshops 2017 (Year: 2017). [cited by applicant]
Sridhar et al. “Interactive Markerless Articulated Hand Motion Tracking Using RGB and Depth Data”, International Conference on Computer Vision (ICCV) 2013 (Year: 2013). [cited by applicant]
Stenger et al. “Model-Based Hand Tracking Using a Hierarchical Bayesian Filter”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, No. 9, Sep. 2006 (Year: 2006). [cited by applicant]
Taylor et al. “User-Specific Hand Modeling from Monocular Depth Sequences”, CVPR2014. (Year: 2014). [cited by applicant]
U.S. Appl. No. 15/432,869, filed Feb. 14, 2017, U.S. Pat. No. 11,841,920, Dec. 12, 2023, Granted. [cited by applicant]
U.S. Appl. No. 15/432,876, filed Feb. 14, 2017, U.S. Pat. No. 11,854,308, Dec. 26, 2023, Granted. [cited by applicant]
U.S. Appl. No. 16/735,440, filed Jan. 6, 2020, Abandoned. [cited by applicant]
U.S. Appl. No. 18/224,373, filed Jul. 20, 2023, Allowed. [cited by applicant]
U.S. Appl. No. 18/536,151, filed Dec. 11, 2023, Pending. [cited by applicant]
U.S. Appl. No. 18/391,574, filed Dec. 20, 2023, Pending. [cited by applicant]