IP Library › Granted Patent US 11,747,912
Granted Patent B1
US 11,747,912 · App. 17/950,825 · Granted Sep 5, 2023

Steerable camera for AR hand tracking

Inventors: Daniel Colascione (Seattle, WA); Patrick Timothy McSweeney Simons (Downey, CA); Weston Welge (Boulder, CO); Ramzi Zahreddine (Denver, CO)
Assignee: Snap Inc.
G06F3/017G02B27/017G06F3/011G06T19/006G02B2027/0138G02B2027/0178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,747,912
App. No.
17/950,825
Granted
Sep 5, 2023
Kind
B1
Abstract

A system for hand tracking for an Augmented Reality (AR) system. The AR system uses a camera of the AR system to capture tracking video frame data of a hand of a user of the AR system. The AR system generates a skeletal model based on the tracking video frame data and determines a location of the hand of the user based on the skeletal model. The AR system causes a steerable camera of the AR system to focus on the hand of the user.

Claims (56)

1. A computer-implemented method comprising:

capturing, by one or more processors, using a first camera of an Augmented Reality (AR) system, tracking video frame data of a hand of a user of the AR system;

generating, by the one or more processors, a skeletal model based on the tracking video frame data;

determining, by the one or more processors, a location of the hand of the user based on the skeletal model;

generating, by the one or more processors, camera steering command data based on the location of the hand of the user; and

focusing, by the one or more processors, a steerable camera of the AR system on the hand of the user.

2. The method of claim 1 , wherein the first camera comprises a non-steerable camera having a camera Field Of View (FOV) equal to an AR FOV of the AR system.

3. The method of claim 1 , wherein the first camera is the steerable camera, the method further comprising:

scanning, by the one or more processors, using the steerable camera, an AR FOV of the AR system to initially locate the hand of the user;

using, by the one or more processors, a look-ahead process to predict a next location of the hand of the user based on a current location of the hand; and

focusing, by the one or more processors, the steerable camera on the next location.

4. The method of claim 3 , wherein the next location is predicted further based on a language model.

5. The method of claim 4 , wherein the language model is for American Sign Language.

6. The method of claim 1 , further wherein the steerable camera comprises:

a camera having sensor and a lens assembly; and

one or more actuators linked to the camera.

7. The method of claim 1 , wherein the AR system comprises a head-worn device.

8. A computing apparatus comprising:

one or more processors; and

a memory storing instructions that, when executed by one or more processors, cause the computing apparatus to perform operations comprising:

capturing, using a first camera of an AR system, tracking video frame data of a hand of a user of the AR system;

generating a skeletal model based on the tracking video frame data;

determining a location of the hand of the user based on the skeletal model;

generating camera steering command data based on the location of the hand of the user; and

focusing a steerable camera of the AR system on the hand of the user.

9. The computing apparatus of claim 8 , wherein the first camera comprises a non-steerable camera having a camera FOV equal to an AR FOV of the AR system.

10. The computing apparatus of claim 8 , wherein the first camera is the steerable camera, and

wherein the instructions, when executed by the one or more processors, further cause the computing apparatus to perform operations comprising:

scanning, using the steerable camera, an AR FOV of the AR system to initially locate the hand of the user;

using a look-ahead process to predict a next location of the hand of the user based on a current location of the hand; and

focusing the steerable camera on the next location.

11. The computing apparatus of claim 10 , wherein the next location is predicted further based on a language model.

12. The computing apparatus of claim 11 , wherein the language model is for American Sign Language.

13. The computing apparatus of claim 8 , wherein the steerable camera comprises:

a camera having sensor and a lens assembly; and

one or more actuators linked to the camera.

14. The computing apparatus of claim 8 , wherein the AR system comprises a head-worn device.

15. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computing apparatus, cause the computing apparatus to perform operations comprising:

capturing, using a first camera of an AR system, tracking video frame data of a hand of a user of the AR system;

generating a skeletal model based on the tracking video frame data;

determining a location of the hand of the user based on the skeletal model;

generating camera steering command data based on the location of the hand of the user; and

focusing a steerable camera of the AR system on the hand of the user.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the first camera comprises a non-steerable camera having a camera FOV equal to an AR FOV of the AR system.

17. The non-transitory computer-readable storage medium of claim 15 , wherein

the first camera is the steerable camera, and

wherein the instructions, when executed by the computing apparatus, further cause the computing apparatus to perform operations comprising:

scanning, using the steerable camera, an AR FOV of the AR system to initially locate the hand of the user;

using a look-ahead process to predict a next location of the hand of the user based on a current location of the hand; and

focusing the steerable camera on the next location.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the next location is predicted further based on a language model.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the language model is for American Sign Language.

20. The non-transitory computer-readable storage medium of claim 15 , wherein the steerable camera comprises:

a camera having sensor and a lens assembly; and

one or more actuators linked to the camera.

21. The non-transitory computer-readable storage medium of claim 15 , wherein the AR system comprises a head-worn device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2023
From: COLASCIONE, DANIEL; SIMONS, PATRICK TIMOTHY MCSWEENEY; WELGE, WESTON; ZAHREDDINE, RAMZI
To: SNAP INC.
Reel/Frame 063878/0986 →
Cited By (3)
US 12,405,675 US 12,535,687 US 12,710,828