IP Library Granted Patent US 12,731,366
Granted Patent B2
US 12,731,366 · App. 18/671,412 · Granted Sep 8, 2026

Computationally adaptive occlusion aware hand landmark estimation

Inventors: Pawan Prasad Bindigan Hariprasanna (Bengaluru, IN); Green Rosh K S (Bengaluru, IN); Vishakha S R (Bengaluru, IN); Sungsoo Choi (Suwon-si, KR); Hyuntaek Woo (Suwon-si, KR); Chaeeun Lee (Suwon-si, KR); Beomsu Kim (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/273G06V10/82G06V40/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,366
App. No.
18/671,412
Granted
Sep 8, 2026
Kind
B2
Abstract

A method performed by an electronic device for estimating a landmark point of a body part of subject by electronic device is provided. The method includes generating, by the electronic device, an initial coarse estimation of the landmark point of the body part using a light-weight deep neural network, determining, by the electronic device, an occluded region of the body part based on the generated initial coarse estimation of the landmark point using a segmentation mask, estimating, by the electronic device, the occlusion probability for the landmark point in the at least one occluded region and the generated initial coarse estimation, determining, by the electronic device, a correction factor for applying on the generated initial coarse estimation as a measure of the estimated occlusion probability, and selecting, by the electronic device, a pre-defined number of neural networks by applying the determined correction factor for processing the at least one occluded region and the generated initial coarse estimation to generate final estimation of the landmark point.

Claims (62)

1 . A method performed by an electronic device for estimating at least one landmark point of a body part of a subject, the method comprising:

generating, by the electronic device, an initial coarse estimation of the at least one landmark point of the body part using a light-weight deep neural network;

determining, by the electronic device, at least one occluded region of the body part based on the generated initial coarse estimation of the at least one landmark point using a segmentation mask;

estimating, by the electronic device, an occlusion probability for the at least one landmark point in the at least one occluded region and the generated initial coarse estimation;

determining, by the electronic device, a correction factor for applying on the generated initial coarse estimation as a measure of the estimated occlusion probability; and

selecting, by the electronic device, a pre-defined number of neural networks by applying the determined correction factor, for processing the at least one occluded region and the generated initial coarse estimation to generate final estimation of the at least one landmark point.

2 . The method of claim 1 , wherein the pre-defined number of the neural networks in each neural network sequence is inversely proportional to the determined correction factor of the at least one landmark point estimation.

3 . The method of claim 1 , wherein light-weight deep neural network predicts a confidence score associated with the generated initial coarse estimation of at least one landmark point of the body part.

4 . The method of claim 1 , further comprising:

generating, by the electronic device, the segmentation mask, wherein the generating of the segmentation mask comprises:

estimating, by the electronic device, a hand bounding box;

performing, by the electronic device, a skin segmentation;

performing, by the electronic device, a connected component analysis for the hand bounding box and the skin segmentation to join at least one component of the body part of the subject; and

generating, by the electronic device, the segmentation mask based on the at least one joined component of the body part of the subject.

5 . The method of claim 1 ,

wherein the occlusion probability for the at least one landmark point is estimated from an occlusion map, and

wherein the occlusion probability differentiates between occlusion due to at least one external object, a self-occlusion by the body part of the subject and a self-occlusion by another body part of the subject.

6 . An electronic device, comprising:

memory storing one or more computer programs;

and

one or more processors communicatively coupled to the memory,

wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors, cause the electronic device to:

generate an initial coarse estimation of at least one landmark point of a body part of a subject using a light-weight deep neural network,

determine at least one occluded region of the body part based on the generated initial coarse estimation of the at least one landmark points using a segmentation mask,

estimate an occlusion probability for the at least one landmark point in the at least one occluded region and the generated initial coarse estimation,

determine a correction factor for applying on the generated initial coarse estimation as a measure of the estimated occlusion probability, and

select a pre-defined number of neural networks by applying the correction factor, for processing the at least one occluded region and the generated initial coarse estimation to generate final estimation of the at least one landmark point.

7 . The electronic device of claim 6 , wherein the pre-defined number of the neural networks in each neural network sequence is inversely proportional to the correction factor of the at least one landmark point estimation.

8 . The electronic device of claim 6 , wherein the light-weight deep neural network predicts a confidence score associated with the generated initial coarse estimation of at least one landmark point of the body part.

9 . The electronic device of claim 6 ,

wherein the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors, cause the electronic device to generate the segmentation mask, and

wherein, to generate the segmentation mask, the one or more computer programs further include computer-executable instructions that, when executed by the one or more processors, cause the electronic device to:

estimate a hand bounding box,

perform a skin segmentation,

perform a connected component analysis for the hand bounding box and the skin segmentation to join at least one component of the body part of the subject, and

generate the segmentation mask based on the at least one joined component of the body part of the subject.

10 . The electronic device of claim 6 ,

wherein the occlusion probability for the at least one landmark point is estimated from an occlusion map, and

wherein the occlusion probability differentiates between occlusion due to at least one external object, a self-occlusion by the body part of the subject and a self-occlusion by another body part of the subject.

11 . The electronic device of claim 8 , wherein the segmentation mask is received together with the confidence score by an occlusion probability estimation engine.

12 . The electronic device of claim 11 , wherein the occlusion probability estimation engine is configured to receive a skin map(S) and a hand mask (H) to determine an external occlusion map (O) equal to an intersection of S and H.

13 . The electronic device of claim 12 , wherein the occlusion probability estimation engine is configured to receive confidence values (C) [1×21] vector and coarse hand landmark (HLM) estimates (Pos) [2×21] vector to determine O[Pos]==0.

14 . The electronic device of claim 13 , wherein, when O[Pos] is zero, the occlusion probability estimation engine is configured to check a confidence value.

15 . The electronic device of claim 14 , wherein, when confidence value is greater than a threshold, the occlusion probability estimation engine is configured to set the occlusion probability to 0.

16 . One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform operations for estimating at least one landmark point of a body part of a subject, the operations comprising:

generating, by the electronic device, an initial coarse estimation of the at least one landmark point of the body part using a light-weight deep neural network;

determining, by the electronic device, at least one occluded region of the body part based on the generated initial coarse estimation of the at least one landmark point using a segmentation mask;

estimating, by the electronic device, an occlusion probability for the at least one landmark point in the at least one occluded region and the generated initial coarse estimation;

determining, by the electronic device, a correction factor for applying on the generated initial coarse estimation as a measure of the estimated occlusion probability; and

selecting, by the electronic device, a pre-defined number of neural networks by applying the determined correction factor, for processing the at least one occluded region and the generated initial coarse estimation to generate final estimation of the at least one landmark point.

17 . The one or more non-transitory computer-readable storage media of claim 16 , wherein the pre-defined number of the neural networks in each neural network sequence is inversely proportional to the correction factor of the at least one landmark point estimation.

18 . The one or more non-transitory computer-readable storage media of claim 16 , wherein light-weight deep neural network predicts a confidence score associated with the generated initial coarse estimation of at least one landmark point of the body part.

19 . The one or more non-transitory computer-readable storage media of claim 16 , the operations further comprising:

generating, by the electronic device, the segmentation mask,

wherein the generating of the segmentation mask comprises:

estimating, by the electronic device, a hand bounding box;

performing, by the electronic device, a skin segmentation;

performing, by the electronic device, a connected component analysis for the hand bounding box and the skin segmentation to join at least one component of the body part of the subject; and

generating, by the electronic device, the segmentation mask based on the at least one joined component of the body part of the subject.

20 . The one or more non-transitory computer-readable storage media of claim 16 ,

wherein the occlusion probability for the at least one landmark point is estimated from an occlusion map, and

wherein the occlusion probability differentiates between occlusion due to at least one external object, a self-occlusion by the body part of the subject and a self-occlusion by another body part of the subject.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2024
From: BINDIGAN HARIPRASANNA, PAWAN PRASAD; ROSH K S, GREEN; S R, VISHAKHA; CHOI, SUNGSOO; WOO, HYUNTAEK; LEE, CHAEEUN; KIM, BEOMSU
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 067494/0755 →
Priority Claims (2)
IN 202341017582 · Mar 15, 2023 · national
IN 2023 41017582 · Feb 13, 2024 · national
Continuity (2)
Continuation PCTKR2024003187 · Mar 12, 2024
Related Publication 20240312176A1 · Sep 19, 2024
References Cited (39)
US 6674877B1 · Jojic · 2004 [cited by examiner]
US 9002099B2 · Litvak et al. · 2015 [cited by applicant]
US 10929654B2 · Iqbal et al. · 2021 [cited by applicant]
US 11783496B2 · Bazarevsky · 2023 [cited by examiner]
US 20090041297A1 · Zhang · 2009 [cited by examiner]
US 20110081053A1 · Zheng · 2011 [cited by examiner]
US 20150110349A1 · Feng · 2015 [cited by examiner]
US 20150347822A1 · Zhou · 2015 [cited by examiner]
US 20180232902A1 · Albadawi · 2018 [cited by examiner]
US 20180365512A1 · Molchanov et al. · 2018 [cited by applicant]
US 20200372246A1 · Chidananda et al. · 2020 [cited by applicant]
US 20200410753A1 · Molyneaux · 2020 [cited by applicant]
US 20210133428A1 · Gernoth · 2021 [cited by examiner]
US 20210174519A1 · Bazarevsky · 2021 [cited by examiner]
US 20210182625A1 · Arar et al. · 2021 [cited by applicant]
US 20220076433A1 · Bazarevsky · 2022 [cited by examiner]
US 20220193888A1 · Rephaeli · 2022 [cited by examiner]
US 20230162472A1 · Gor · 2023 [cited by examiner]
US 20250086820A1 · Raitarovskyi · 2025 [cited by examiner]
US 20250292616A1 · Lin · 2025 [cited by examiner]
CN 111027407A · 2020 [cited by applicant]
CN 111191632A · 2020 [cited by applicant]
CN 115131843A · 2022 [cited by applicant]
IN 202011043049 · 2020 [cited by applicant]
KR 102305403B1 · 2021 [cited by applicant]
KR 1020220164376A · 2022 [cited by applicant]
WO WO2025042229A1 · 2025 [cited by examiner]
Gu Y, Zhang H, Kamijo S. Multi-Person Pose Estimation Using an Orientation and Occlusion Aware Deep Learning Network. Sensors (Basel). Mar. 12, 2020;20(6):1593. doi: 10.3390/s20061593. PMID: 32178461; PMCID: PMC7146407.… [cited by examiner]
G. Yan, Z. Wang, S. Geng, Y. Yu and Y. Guo, “Part-Based Representation Enhancement for Occluded Person Re-Identification,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, No. 8, pp. 4217-4231… [cited by examiner]
B. Chen, T.-J. Chin and M. Klimavicius, “Occlusion-Robust Object Pose Estimation with Holistic Representation,” 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2022, pp. 222… [cited by examiner]
C. Cimarelli, et al, “A case study on the impact of masking moving objects on the camera pose regression with CNNs,” 2019 16th IEEE Intl Conf on Advanced Video and Signal Based Surveillance (AVSS), Taipei, Taiwan, 2019,… [cited by examiner]
L. Ke, M.-C. Chang, H. Qi and S. Lyu, “DetPoseNet: Improving Multi-Person Pose Estimation via Coarse-Pose Filtering,” in IEEE Transactions on Image Processing, vol. 31, pp. 2782-2795, 2022, doi: 10.1109/TIP.2022.3161081… [cited by examiner]
Mohamed S. Abdallah et al., “Light-Weight Deep Learning Techniques with Advanced Processing for Real-Time Hand Gesture Recognition”, Sensors 2023, vol. 23, No. 1, pp. 1-20, Dec. 20, 2022. [cited by applicant]
International Search Report and Written Opinion dated Jun. 21, 2024, issued in International Application No. PCT/KR2024/003187. [cited by applicant]
Indian Office Action dated Jul. 8, 2025, issued in an Indian Patent Application No. 202341017582. [cited by applicant]
European Search Report dated Jan. 2, 2026, issued in European Application No. 24771174.0. [cited by applicant]
Fan Zhang et al., MediaPipe Hands: On-device Real-time Hand Tracking, XP081698245, Jun. 18, 2020, Ithaca, New York. [cited by applicant]
Meng et al., 3D Interacting Hand Pose Estimation by Hand De-occlusion and Removal, ECCV 2022. [cited by applicant]
Park et al., HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation Network, CVPR 2022. [cited by applicant]