IP Library › Granted Patent US 12,260,567
Granted Patent B2
US 12,260,567 · App. 17/973,167 · Granted Mar 25, 2025

3D space carving using hands for object capture

Inventors: Branislav Micusik (St. Andrae-Woerdern, AT); Georgios Evangelidis (Vienna, AT); Daniel Wolf (Mödling, AT)
Assignee: Snap Inc.
G06T7/292G06F3/011G06T7/564G06T19/006G06V20/64G06T2207/10012G06T2207/10028G06T2210/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,567
App. No.
17/973,167
Granted
Mar 25, 2025
Kind
B2
Abstract

A method for carving a 3D space using hands tracking is described. In one aspect, a method includes accessing a first frame from a camera of a display device, tracking, using a hand tracking algorithm operating at the display device, hand pixels corresponding to one or more user hands depicted in the first frame, detecting, using a sensor of the display device, depths of the hand pixels, identifying a 3D region based on the depths of the hand pixels, and applying a 3D reconstruction engine to the 3D region.

Claims (66)

1. A method comprising:

accessing a first frame from a camera of a display device;

tracking, using a hand tracking algorithm operating at the display device, hand pixels corresponding to one or more user hands depicted in the first frame;

detecting, using a sensor of the display device, depths of the hand pixels;

tracking a motion of a left hand and a right hand of a user of the display device;

carving a 3D region that includes a 3D hull of a physical object based on the motion of the left hand and the right hand relative to the physical object;

identifying the 3D region based on the depths of the hand pixels; and

applying a 3D reconstruction engine to the 3D region.

2. The method of claim 1 , wherein the 3D region includes an unoccupied 3D space between the camera and the one or more user hands.

3. The method of claim 1 , wherein the sensor includes a depth sensor or stereo cameras.

4. The method of claim 1 , wherein detecting the depths is based on contour matching of the one or more user hands in two images.

5. The method of claim 1 , wherein carving the 3D region comprises:

detecting the 3D region between the left hand and right hand; and

identifying a 3D envelope comprising the physical object based on the motion of the left hand and the right hand, the motion of the left hand indicating a left side boundary of the 3D region, the right hand indicating a right side boundary of the 3D region.

6. The method of claim 5 , wherein applying the 3D reconstruction engine to the 3D region comprises:

generating a 3D model of the physical object included in the 3D envelope based on point cloud data from the 3D envelope.

7. The method of claim 6 , further comprising:

identifying the physical object based on the 3D model of the physical object.

8. The method of claim 7 , further comprising:

identifying virtual content corresponding to the physical object or the 3D model of the physical object; and

displaying, in a display of the display device, the virtual content as an overlay to the physical object.

9. The method of claim 1 , wherein identifying the 3D region is based on a motion of the one or more user hands comprises:

filtering a first portion of the first frame to identify a first area of interest based on a location of the one or more user hands in the first frame;

filtering a second portion of a second frame to identify a second area of interest based on a location of the one or more user hands in the second frame;

identifying first hand pixel depths of the one or more user hands in the first frame;

identifying second hand pixel depths of the one or more user hands in the second frame; and

identifying the 3D region based on the first area of interest, the second area of interest, the first hand pixel depths, and the second hand pixel depths.

10. The method of claim 1 , wherein carving the 3D region comprises:

detecting the 3D region based on the motion of the left hand and the right hand, the left hand and right hand moving in front of the physical object, behind the physical object, and adjacent to the physical object.

11. A computing apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the apparatus to:

access a first frame from a camera of a display device;

track, using a hand tracking algorithm operating at the display device, hand pixels corresponding to one or more user hands depicted in the first frame;

detect, using a sensor of the display device, depths of the hand pixels;

track a motion of a left hand and a right hand of a user of the display device;

carve a 3D region that includes a 3D hull of a physical object based on the motion of the left hand and the right hand relative to the physical object;

identify the 3D region based on the depths of the hand pixels; and

apply a 3D reconstruction engine to the 3D region.

12. The computing apparatus of claim 11 , wherein the 3D region includes an unoccupied 3D space between the camera and the one or more user hands.

13. The computing apparatus of claim 11 , wherein the sensor includes a depth sensor or stereo cameras.

14. The computing apparatus of claim 11 , wherein detecting the depths is based on contour matching of the one or more user hands in two images.

15. The computing apparatus of claim 11 , wherein carving the 3D region comprises:

detecting the 3D region between the left hand and right hand tracking; and

identifying a 3D envelope comprising the physical object based on the motion of the left hand and the right hand, the motion of the left hand indicating a left side boundary of the 3D region, the right hand indicating a right side boundary of the 3D region.

16. The computing apparatus of claim 15 , wherein applying the 3D reconstruction engine to the 3D region comprises:

generate a 3D model of the physical object included in the 3D envelope based on point cloud data from the 3D envelope.

17. The computing apparatus of claim 16 , wherein the instructions further configure the apparatus to:

identify the physical object based on the 3D model of the physical object.

18. The computing apparatus of claim 17 , wherein the instructions further configure the apparatus to:

identify virtual content corresponding to the physical object or the 3D model of the physical object; and

display, in a display of the display device, the virtual content as an overlay to the physical object.

19. The computing apparatus of claim 11 , wherein identifying the 3D region is based on a motion of the one or more user hands comprises:

filter a first portion of the first frame to identify a first area of interest based on a location of the one or more user hands in the first frame;

filter a second portion of a second frame to identify a second area of interest based on a location of the one or more user hands in the second frame;

identify first hand pixel depths of the one or more user hands in the first frame;

identify second hand pixel depths of the one or more user hands in the second frame; and

identify the 3D region based on the first area of interest, the second area of interest, the first hand pixel depths, and the second hand pixel depths.

20. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to:

access a first frame from a camera of a display device;

track, using a hand tracking algorithm operating at the display device, hand pixels corresponding to one or more user hands depicted in the first frame;

detect, using a sensor of the display device, depths of the hand pixels;

track a motion of a left hand and a right hand of a user of the display device;

carve a 3D region that includes a 3D hull of a physical object based on the motion of the left hand and the right hand relative to the physical object;

identify the 3D region based on the depths of the hand pixels; and

apply a 3D reconstruction engine to the 3D region.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2022
From: MICUSIK, BRANISLAV; EVANGELIDIS, GEORGIOS; WOLF, DANIEL
To: SNAP INC.
Reel/Frame 061531/0493 →
Priority Claims (1)
GR 20220100720 · Sep 1, 2022 · national
Continuity (2)
Related Publication 20240135555A1 · Apr 25, 2024
Related Publication 20240233144A9 · Jul 11, 2024
References Cited (16)
US 8073198B2 · Marti · 2011 [cited by applicant]
US 9639943B1 · Kutliroff · 2017 [cited by examiner]
US 10706584B1 · Ye · 2020 [cited by examiner]
US 20170242492A1 · Horowitz · 2017 [cited by examiner]
US 20180173404A1 · Smith · 2018 [cited by examiner]
EP 2956843 · 2019 [cited by applicant]
WO 2022040954 · 2022 [cited by applicant]
WO WO2022040954A1 · 2022 [cited by examiner]
Kutulakos, Kiriakos N., and James R. ValliNo. “Calibration-free augmented reality.” IEEE Transactions on Visualization and Computer Graphics 4, No. 1 (1998): 1-20. (Year: 1998). [cited by examiner]
Michel, Damien, Xenophon Zabulis, and Antonis A. Argyros. “Shape from interaction.” Machine Vision and Applications 25 (2014): 1077-1087. (Year: 2014). [cited by examiner]
Tzionas, Dimitrios, and Juergen Gall. “3d object reconstruction from hand-object interactions.” In Proceedings of the IEEE International Conference on Computer Vision, pp. 729-737. 2015. (Year: 2015). [cited by examiner]
Ye Y, Gupta A, Tulsiani S. What's in your hands? 3D Reconstruction of Generic Objects in Hands. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 18, 2022 (pp. 3885-3895). IEEE. (Year: 20… [cited by examiner]
Machine translation of WO 2022/040954 (Year: 2022). [cited by examiner]
“International Application Serial No. PCT/US2023/073217, International Search Report mailed Dec. 13, 2023”, 5 pgs. [cited by applicant]
“International Application Serial No. PCT/US2023/073217, Written Opinion mailed Dec. 13, 2023”, 11 pgs. [cited by applicant]
Kim, N H, “A Contour-Based Stereo Matching Algorithm Using Disparity Continuity”, Pattern Recognition, Elsevier, GB, vol. 21, No. 5, (Jan. 1, 1988), 505-514. [cited by applicant]