IP Library Granted Patent US 12,530,087
Granted Patent B1
US 12,530,087 · App. 18/357,090 · Granted Jan 20, 2026

Tracking interacting hands using sensor of wearable multimedia device

Inventors: Sylvana Alpert (Sunnyvale, CA); Ralph Brunner (Los Gatos, CA)
Assignee: Hewlett-Packard Development Company, L.P.
G06F3/017G01S17/894G06F3/011G06V40/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,087
App. No.
18/357,090
Granted
Jan 20, 2026
Kind
B1
Abstract

Systems, methods, devices and non-transitory, computer-readable storage mediums are disclosed for the tracking of interacting hands using a sensor of a wearable multimedia device. In an embodiment, a method comprises: obtaining, with a sensor of a wearable multimedia device, a first frame of two-dimensional (2D) image data and a second frame of three-dimensional (3D) depth data; determining multiple hand regions in the first frame of 2D image data; determining a location of each hand region in the first frame of 2D image data; detecting at least one landmark in each detected hand region; generating a confidence score for each landmark detected in each detected hand region; determining 3D world coordinates for each landmark that is visible in the first frame of the 2D image data; tracking each landmark in 3D world coordinates; and determining an interaction with the hands based on the tracking of each landmark in 3D world coordinates.

Claims (41)

1 . A method comprising:

obtaining, with a sensor of a wearable multimedia device, a first frame of two-dimensional (2D) image data and a second frame of three-dimensional (3D) depth data;

determining, with at least one processor, multiple hand regions in the first frame of 2D image data;

determining, with the at least one processor, a location of each hand region in the first frame of 2D image data;

detecting, with the at least one processor, at least one landmark in each detected hand region;

generating, with the at least one processor, a confidence score for each landmark detected in each detected hand region;

determining, with the at least one processor, 3D world coordinates for each landmark that is visible in the first frame of the 2D image data;

tracking, with the at least one processor, each landmark in 3D world coordinates; and

determining, with the at least one processor, an interaction with the hands based on the tracking of each landmark in 3D world coordinates.

2 . The method of claim 1 , wherein the at least one landmark is a joint or finger tip of a finger in the detected hand region.

3 . The method of claim 1 , wherein the at least one landmark is a palm or wrist in the detected hand region.

4 . The method of claim 1 , wherein detecting at least one landmark in each detected hand region comprises using a landmark detection model that is trained to predict the 2D or 3D point coordinates of the at least one landmark.

5 . The method of claim 1 , wherein bounding boxes for the location of each hand region in the first frame of 2D image data and for the at least one landmark in each detected hand region are predicted by a machine learning model, and the bounding boxes identify locations of the hand regions and landmarks in the 2D image data.

6 . The method of claim 5 , wherein the landmark detection model is trained on annotated ground truth data and a synthetic hand model over various backgrounds that is mapped to corresponding 2D or 3D point coordinates.

7 . The method of claim 1 , wherein the sensor is a time of flight camera that outputs infrared or amplitude image data and the depth data that are registered to each other by the sensor.

8 . The method of claim 7 , wherein determining 3D world coordinates in the second frame of 3D depth data for each landmark that is visible in the first frame of the 2D image data comprises adding a depth component from the depth data to the corresponding 2D pixel coordinates of each landmark.

9 . The method of claim 7 , wherein the time of flight camera is an infrared camera that is adjusted to measure a range of temperature that approximates human body temperature, and each hand region is detected by binarization on the first frame using a threshold value.

10 . The method of claim 1 , wherein each hand region is detected using template matching.

11 . The method of claim 1 , wherein the landmark is determined to be visible based at least in part on the confidence score for the landmark.

12 . A system comprising:

a sensor;

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

obtaining, with the sensor, a first frame of two-dimensional (2D) image data and a second frame of three-dimensional (3D) depth data;

determining multiple hand regions in the first frame of 2D image data;

determining a location of each hand region in the first frame of 2D image data;

detecting at least one landmark in each detected hand region;

generating a confidence score for each landmark detected in each detected hand region;

determining 3D world coordinates for each landmark that is visible in the first frame of the 2D image data;

tracking each hand in 3D world coordinates; and

determining an interaction with the hands based on the tracking of each landmark in 3D world coordinates.

13 . The system of claim 12 , wherein the at least one landmark is a joint or finger tip of a finger in the detected hand region.

14 . The system of claim 12 , wherein the at least one landmark is a palm in the detected hand region.

15 . The system of claim 12 , wherein detecting at least one landmark in each detected hand region comprises using a landmark detection model that is trained to predict the 2D or 3D world coordinates of the at least one landmark.

16 . The system of claim 12 , wherein bounding boxes for the location of each hand region in the first frame of 2D image data and for the at least one landmark in each detected hand region are predicted by a machine learning model, and the bounding boxes identify locations of the hand regions and landmarks in the 2D image data.

17 . The system of claim 16 , wherein the landmark detection model is trained on annotated ground truth data and a synthetic hand model over various backgrounds that is mapped to corresponding 3D world coordinates.

18 . The system of claim 17 , wherein determining 3D world coordinates in the second frame of 3D depth data for each landmark that is visible in the first frame of the 2D image data comprises adding a depth component from the depth data to the corresponding 2D pixel coordinates of each landmark.

19 . The system of claim 12 , wherein the sensor is a time of flight camera that outputs infrared or amplitude image data and the depth data that are registered by the sensor.

20 . The system of claim 19 , wherein the time of flight camera is an infrared camera that is adjusted to measure a range of temperature that approximates human body temperature, and each hand region is detected by binarization on the first frame using a threshold value.

21 . The system of claim 12 , wherein each hand region is detected using template matching.

22 . The system of claim 12 , wherein the landmark is determined to be visible based at least in part on the confidence score for the landmark.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: HUMANE, INC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 071844/0747 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2025
From: ALPERT, SYLVANA; BRUNNER, RALPH
To: HUMANE, INC.
Reel/Frame 070262/0025 →
Continuity (1)
Provisional Application 63394590 · Aug 2, 2022
References Cited (14)
US 10353532B1 · Holz · 2019 [cited by examiner]
US 20200334828A1 · Öztireli · 2020 [cited by examiner]
US 20210174519A1 · Bazarevsky · 2021 [cited by examiner]
US 20230298283A1 · Despande · 2023 [cited by examiner]
US 20240331446A1 · Gao · 2024 [cited by examiner]
Kanel, “Sixth Sense Technology,” Thesis for the Bachelor Degree of Engineering in Information and Technology, Centria University of Applied Sciences, May 2014, 46 pages. [cited by applicant]
Mann et al., “Telepointer: Hands-Free Completely Self Contained Wearable Visual Augmented Reality without Headwear and without any Infrastructural Reliance”, IEEE International Symposium on Wearable Computing, Oct. 2000… [cited by applicant]
Mann, “Wearable Computing: a First Step Toward Personal Imaging,” IEEE Computer, Feb. 1997, 30(2):25-32. [cited by applicant]
Mann, “Wearable, tetherless computer-mediated reality,” Feb. 1996. In Presentation at the American Association of Artificial Intelligence, 1996 Symposium; early draft appears as MIT Media Lab Technical Report 260, Dec. … [cited by applicant]
Metavision.com [online], “Sensularity with a Sixth Sense,” available on or before Apr. 7, 2015, via Internet Archive: Wayback Machine URL <http://web.archive.org/web/20170901072037/https://blog.metavision.com/professor-… [cited by applicant]
Microsoft.com [online], “Direct manipulation with hands,” Sep. 2022, retrieved on Jul. 12, 2023, retrieved from URL<https://learn.microsoft.com/en-us/windows/mixed-reality/design/direct-manipulation>, 21 pages. [cited by applicant]
Microsoft.com [online], “Hand tracking—MRTK2,” Aug. 2022, retrieved on Jul. 12, 2023, retrieved from URL<https://learn.microsoft.com/en-us/windows/mixed-reality/mrtk-unity/mrtk2/features/input/hand-tracking?view=mrtkuni… [cited by applicant]
Mistry et al., “WUW—wear Ur world: a wearable gestural interface”, Proceedings of CHI EA '09 Extended Abstracts on Human Factors in Computing Systems, ACM New York, NY, USA, 6 pages. [cited by applicant]
Shetty et al., “Sixth Sense Technology,” International Journal of Science and Research, Dec. 2014, 3(12):1068-1073. [cited by applicant]