IP Library › Granted Patent US 11,551,374
Granted Patent B2
US 11,551,374 · App. 17/015,819 · Granted Jan 10, 2023

Hand pose estimation from stereo cameras

Inventors: Yuncheng Li (Los Angeles, CA); Jonathan M. Rodriguez, II (Los Angeles, CA); Zehao Xue (Los Angeles, CA); Yingying Wang (Marina del Rey, CA)
Assignee: Snap Inc.
G06T7/73G06F3/011G06K9/6256G06T7/20G06V40/11H04N13/204G06T2207/10012G06T2207/20081G06T2207/20084G06T2207/20132G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,374
App. No.
17/015,819
Granted
Jan 10, 2023
Kind
B2
Abstract

Systems and methods herein describe using a neural network to identify a first set of joint location coordinates and a second set of joint location coordinates and identifying a three-dimensional hand pose based on both the first and second sets of joint location coordinates.

Claims (59)

1. A method comprising:

receiving, from a camera, a plurality of images of a hand;

generating a plurality of sets of joint location coordinates by:

for each given image in the plurality of images:

cropping, using one or more processors, a portion of the given image comprising the hand;

identifying, using a neural network, a first set of joint location coordinates in the cropped portion of the given image, wherein the first set of joint location coordinates represent pixel locations of the hand relative to the cropped portion of the given image;

identifying an intermediate set of joint location coordinates, wherein the intermediate set of joint location coordinates represent pixel locations of the hand relative to the given image;

generating a second set of joint location coordinates based on the first set of joint location coordinates and the intermediate set, wherein the second set of joint location coordinates represents joint locations of the hand relative to a three-dimensional physical space; and

identifying a three-dimensional hand pose of the hand based on the plurality of sets of joint location coordinates.

2. The method of claim 1 , wherein the plurality of images comprises a plurality of views of the hand.

3. The method of claim 1 , further comprising:

prompting a user of a client device to initialize a hand position;

receiving the initialized hand position; and

tracking the hand based on the initialized hand position.

4. The method of claim 1 , wherein the camera is a stereo camera.

5. The method of claim 1 ,

wherein the intermediate set of joint location coordinates is measured relative to an uncropped version of the given image.

6. The method of claim 1 , further comprising:

generating a synthetic training dataset comprising stereo image pairs of virtual hands and corresponding ground truth labels, wherein the corresponding ground truth labels comprise joint locations.

7. A system comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the system to perform operations comprising:

receiving, from a camera, a plurality of images of a hand;

generating a plurality of sets of joint location coordinates by:

for each given image in the plurality of images:

cropping, using one or more processors, a portion of the given image comprising the hand;

identifying, using a neural network, a first set of joint location coordinates in the cropped portion of the given image, wherein the first set of joint location coordinates represent pixel locations of the hand relative to the cropped portion of the given image;

identifying an intermediate set of joint location coordinates, wherein the intermediate set of joint location coordinates represent pixel locations of the hand relative to the given image;

generating a second set of joint location coordinates based on the first set of joint location coordinates and the intermediate set, wherein the second set of joint location coordinates represents joint locations of the hand relative to a three-dimensional physical space; and

identifying a three-dimensional hand pose of the hand based on the plurality of sets of joint location coordinates.

8. The system of claim 7 , wherein the plurality of images comprises a plurality of views of the hand.

9. The system of claim 7 , further comprising:

prompting a user of a client device to initialize a hand position;

receiving the initialized hand position; and

tracking the hand based on the initialized hand position.

10. The system of claim 7 , wherein the camera is a stereo camera.

11. The system of claim 7 ,

wherein the intermediate set of joint location coordinates is measured relative to an uncropped version of the given image.

12. The system of claim 7 , further comprising:

generating a synthetic training dataset comprising stereo image pairs of virtual hands and corresponding ground truth labels, wherein the corresponding ground truth labels comprise joint locations.

13. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to perform operations comprising:

receiving, from a camera, a plurality of images of a hand;

generating a plurality of sets of joint location coordinates by:

for each given image in the plurality of images:

cropping, using one or more processors, a portion of the given image comprising the hand;

identifying, using a neural network, a first set of joint location coordinates in the cropped portion of the given image, wherein the first set of joint location coordinates represent pixel locations of the hand relative to the cropped portion of the given image;

identifying an intermediate set of joint location coordinates, wherein the intermediate set of joint location coordinates represent pixel locations of the hand relative to the given image;

generating a second set of joint location coordinates based on the first set of joint location coordinates and the intermediate set, wherein the second set of joint location coordinates represents joint locations of the hand relative to a three-dimensional physical space; and

identifying a three-dimensional hand pose of the hand based on the plurality of sets of joint location coordinates.

14. The computer-readable storage medium of claim 13 , wherein the plurality of images comprises a plurality of views of the hand.

15. The computer-readable storage medium of claim 13 , further comprising:

prompting a user of a client device to initialize a hand position;

receiving the initialized hand position; and

tracking the hand based on the initialized hand position.

16. The computer-readable storage medium of claim 13 , wherein the second set of joint location coordinates is measured use millimeters.

17. The computer-readable storage medium of claim 13 , further comprising:

generating a synthetic training dataset comprising stereo image pairs of virtual hands and corresponding ground truth labels, wherein the corresponding ground truth labels comprise joint locations.

18. The computer-readable storage medium of claim 13 ,

wherein the intermediate set of joint location coordinates is measured relative to an uncropped version of the given image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2022
From: LI, YUNCHENG; RODRIGUEZ, JONATHAN M, II; XUE, ZEHAO; WANG, YINGYING
To: SNAP INC.
Reel/Frame 061636/0273 →
Continuity (2)
Provisional Application 62897669 · Sep 9, 2019
Related Publication 20210074016A1 · Mar 11, 2021
Cited By (5)
US 12,340,627 US 12,366,920 US 12,366,923 US 12,502,110 US 12,572,220