IP Library › Granted Patent US 12,573,080
Granted Patent B2
US 12,573,080 · App. 18/082,190 · Granted Mar 10, 2026

Electronic device and method for obtaining three-dimensional (3D) skeleton data of object captured using plurality of cameras

Inventors: Deokho Kim (Suwon-si, KR); Taehyuk Kwon (Suwon-si, KR); Hwangpil Park (Suwon-si, KR); Jiwon Jeong (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T7/73G06V10/25G06V10/70G06V40/11G06T2207/10012G06T2207/20072G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,080
App. No.
18/082,190
Granted
Mar 10, 2026
Kind
B2
Abstract

An electronic device performs a method of obtaining Three-Dimensional (3D) skeleton data of an object obtained by using a first camera and a second camera. The method includes: obtaining a first image using the first camera and obtaining a second image using the second camera; obtaining, from the first image, a first Region Of Interest (ROI) comprising the object; obtaining, from the first ROI, first skeleton data comprising at least one keypoint of the object; obtaining a second ROI from the second image, based on the first skeleton data and information about a relative position between the first camera and the second camera; obtaining, from the second ROI, second skeleton data comprising at least one keypoint of the object; and obtaining 3D skeleton data of the object, based on the first skeleton data and the second skeleton data.

Claims (36)

1 . A method, performed by an electronic device, of obtaining Three-Dimensional (3D) skeleton data of an object by using a first camera and a second camera, the method comprising:

obtaining a first image using the first camera and obtaining a second image using the second camera;

obtaining, from the first image, a first Region Of Interest (ROI) within the first image, the first ROI comprising the object;

obtaining, from the first ROI within the first image, first skeleton data comprising at least one keypoint of the object;

projecting 3D position coordinates of the at least one keypoint of the object to two-dimensional (2D) position coordinates on the second image, based on information about a relative position between the first camera and the second camera;

identifying, in the second image, second ROI having the 2D position coordinates of the at least one keypoint;

obtaining, from the second ROI, second skeleton data comprising the at least one keypoint of the object; and

obtaining the 3D skeleton data of the object, based on the first skeleton data and the second skeleton data.

2 . The method of claim 1 , wherein the object is a hand or a body part of a user.

3 . The method of claim 1 , wherein, in the obtaining of the first ROI from the first image, a first deep learning model trained to use the first image as an input value and output an image region comprising the object as the first ROI is used.

4 . The method of claim 1 , wherein, in the obtaining of the first skeleton data from the first ROI, a second deep learning model trained to use the first ROI as an input value and output, as the first skeleton data, a graph comprising the at least one keypoint of the object as a joint is used.

5 . The method of claim 1 , wherein the first skeleton data comprises 3D position coordinates of the at least one keypoint with respect to a preset origin in space.

6 . The method of claim 1 , wherein the second skeleton data comprises 2D position coordinates of the at least one keypoint with respect to a preset origin on a plane of the second image.

7 . The method of claim 6 , wherein, in the obtaining of the second skeleton data from the second ROI, a third deep learning model trained to use the second ROI as an input value and output, as the second skeleton data, a graph comprising the at least one keypoint of the object as a joint is used.

8 . The method of claim 6 , wherein the first skeleton data comprises 3D position coordinates of the at least one keypoint with respect to a preset origin in space, and

wherein the obtaining of the 3D skeleton data of the object based on the first skeleton data and the second skeleton data comprises:

projecting the 3D position coordinates of the at least one keypoint, the 3D position coordinates being included in the first skeleton data, to 2D position coordinates on a plane of the first image; and

obtaining 3D position coordinates, based on the 2D position coordinates projected on the plane of the first image and the 2D position coordinates included in the second skeleton data.

9 . The method of claim 1 , wherein the second skeleton data comprises 3D position coordinates of the at least one keypoint with respect to a preset origin in space.

10 . The method of claim 9 , wherein, in the obtaining of the second skeleton data from the second ROI, a second deep learning model trained to use the second ROI as an input value and output, as the second skeleton data, a graph comprising the at least one keypoint of the object as a joint is used.

11 . The method of claim 9 , wherein the first skeleton data comprises 3D position coordinates of the at least one keypoint with respect to a preset origin in space, and

wherein the obtaining of the 3D skeleton data of the object based on the first skeleton data and the second skeleton data comprises obtaining 3D position coordinates of the at least one keypoint, based on an average value of the 3D position coordinates included in the first skeleton data and the 3D position coordinates included in the second skeleton data.

12 . The method of claim 1 , wherein the obtaining of the 3D skeleton data of the object based on the first skeleton data and the second skeleton data comprises obtaining 3D position coordinates of the at least one keypoint, based on a value obtained by weight-combining position coordinates in the first skeleton data with position coordinates in the second skeleton data.

13 . An electronic device comprising:

a first camera;

a second camera;

a storage storing at least one instruction; and

at least one processor configured to electrically connect with the first camera and the second camera and configured to execute the at least one instruction stored in the storage to:

obtain a first image via the first camera and obtain a second image via the second camera,

obtain, from the first image, a first Region Of Interest (ROI) within the first image, the first ROI comprising an object,

obtain, from the first ROI within the first image, first skeleton data comprising at least one keypoint of the object,

projecting three-dimensional (3D) position coordinates of the at least one keypoint of the object to two-dimensional (2D) position coordinates on the second image, based on information about a relative position between the first camera and the second camera,

identifying, in the second image, second ROI having the 2D position coordinates of the at least one keypoint,

obtain, from the second ROI, second skeleton data comprising the at least one keypoint of the object, and

obtain the 3D skeleton data of the object, based on the first skeleton data and the second skeleton data.

14 . A non-transitory computer-readable recording medium having recorded thereon a program for performing the method of claim 1 , on a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2022
From: KIM, DEOKHO; KWON, TAEHYUK; PARK, HWANGPIL; JEONG, JIWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062109/0373 →
Priority Claims (1)
KR 10-2021-0179973 · Dec 15, 2021 · national
Continuity (2)
Continuation PCTKR2022020374 · Dec 14, 2022
Related Publication 20230186512A1 · Jun 15, 2023
References Cited (32)
US 9877012B2 · Uchiyama et al. · 2018 [cited by applicant]
US 10495450B2 · Nakagawa et al. · 2019 [cited by applicant]
US 10552971B2 · Gu et al. · 2020 [cited by applicant]
US 10853958B2 · Choi · 2020 [cited by applicant]
US 10937237B1 · Kim · 2021 [cited by examiner]
US 11386637B2 · Park et al. · 2022 [cited by applicant]
US 11532127B2 · Meng et al. · 2022 [cited by applicant]
US 11538207B2 · Li et al. · 2022 [cited by applicant]
US 11798177B2 · Wu · 2023 [cited by applicant]
US 20120106784A1 · Cho et al. · 2012 [cited by applicant]
US 20150243072A1 · Haglund · 2015 [cited by examiner]
US 20180024641A1 · Mao et al. · 2018 [cited by applicant]
US 20210074016A1 · Li · 2021 [cited by examiner]
US 20210174519A1 · Bazarevsky · 2021 [cited by examiner]
US 20230009367A1 · Goodman · 2023 [cited by examiner]
US 20230028562A1 · Papon · 2023 [cited by examiner]
US 20230298204A1 · Wang · 2023 [cited by examiner]
CN 112927259A · 2021 [cited by applicant]
EP 3477543A1 · 2019 [cited by applicant]
JP 201932600A · 2019 [cited by applicant]
JP 202156922A · 2021 [cited by applicant]
KR 1020120044484A · 2012 [cited by applicant]
KR 101711736B1 · 2017 [cited by applicant]
KR 102138680B1 · 2020 [cited by applicant]
KR 1020210009458A · 2021 [cited by applicant]
KR 1020210011425A · 2021 [cited by applicant]
KR 1020210036879A · 2021 [cited by applicant]
KR 1020210150881A · 2021 [cited by applicant]
Communications dated Mar. 20, 2023, issued by the International Searching Authority in counterpart International Application No. PCT/KR2022/020374 (PCT/ISA/220, PCT/ISA/210, and PCT/ISA/237). [cited by applicant]
Simon et al., “Hand Keypoint Detection in Single Images Using Multiview Bootstrapping”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), ArXiv, Apr. 25, 2017, 9 total pages, arXiv:1704.07809v1, do… [cited by applicant]
Communication dated Nov. 13, 2024, issued by European Patent Office in European Patent Application No. 22907938.9. [cited by applicant]
Communication dated Jul. 11, 2025, issued by the European Patent Office in counterpart European Application No. 22907938.9. [cited by applicant]