IP Library Granted Patent US 12694613
Granted Patent B2
US 12694613 · App. 18/655,969 · Granted Jul 28, 2026

Method, apparatus and electronic device for hand three-dimensional reconstruction

Inventors: Xiaozheng Zheng (Beijing, CN); Chao Wen (Beijing, CN); Zhou Xue (Beijing, CN)
Assignee: Beijing Zitiao Network Technology Co., Ltd.
G06T17/00G06T19/20G06V10/7715G06V10/776G06V10/806G06V10/82G06V40/107G06T2219/2016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694613
App. No.
18/655,969
Granted
Jul 28, 2026
Kind
B2
Abstract

The embodiments of the disclosure provide a method, apparatus, and electronic device for hand three-dimensional reconstruction. One specific implementation of the method includes: obtaining hand images acquired at at least two angles of view; determining an initial hand three-dimensional reconstruction result corresponding to each of the hand images based on a predetermined hand three-dimensional reconstruction network, wherein the hand three-dimensional reconstruction result comprises a hand three-dimensional model and a hand key point; and fusing the initial hand three-dimensional reconstruction results corresponding to the hand images acquired at the at least two angles of view to obtain a fused hand three-dimensional reconstruction result. With consistency of a plurality of angles of views, the implementation may make the fused hand three-dimensional reconstruction result more accurate.

Claims (71)

1 . A method of hand three-dimensional reconstruction, implemented using a dedicated hardware-based system with a processor, comprising:

obtaining hand images acquired at at least two angles of view;

determining, by using an interaction feature, an initial hand three-dimensional reconstruction result corresponding to each of the hand images based on a predetermined hand three-dimensional reconstruction network, wherein the interaction feature is obtained by performing interaction on hand features corresponding to each of the hand images, and the initial hand three-dimensional reconstruction result comprises a hand three-dimensional model and a hand key point; and

fusing the initial hand three-dimensional reconstruction results corresponding to the hand images acquired at the at least two angles of view to obtain a fused hand three-dimensional reconstruction result.

2 . The method of claim 1 , wherein determining, by using the interaction feature, the initial hand three-dimensional reconstruction result corresponding to each of the hand images based on the predetermined hand three-dimensional reconstruction network comprises:

for each hand image of the hand images,

determining a hand feature corresponding to the hand image based on the predetermined hand three-dimensional reconstruction network;

updating the hand feature corresponding to the hand image by using the interaction feature to obtain an updated hand feature corresponding to the hand image; and

determining the initial hand three-dimensional reconstruction result corresponding to the hand image based on the hand three-dimensional reconstruction network and the updated hand feature corresponding to the hand image.

3 . The method of claim 2 , wherein determining the hand feature corresponding to the hand image based on the predetermined hand three-dimensional reconstruction network comprises:

determining target information by inputting the hand image into the predetermined hand three-dimensional reconstruction network, wherein the target information comprises a feature map of at least one level and the hand key point; and

determining the hand feature corresponding to the hand image based on the target information.

4 . The method of claim 3 , wherein the target information comprises a hand gesture vector; and

determining the hand feature corresponding to the hand image based on the target information comprising:

determining the hand feature corresponding to the hand image based on at least one of a first hand feature, a second hand feature and a third hand feature; and wherein

the first hand feature is obtained by encoding the hand gesture vector and a coordinate of the hand key point;

the second hand feature is determined by processing a feature map of a target level using a predefined graph algorithm; and

the third hand feature is obtained by projecting the hand key point onto feature maps of levels other than the feature map of the target level, and determining features at projection position points on respective feature maps.

5 . The method of claim 4 , wherein determining the hand feature corresponding to the hand image based on at least one of the first hand feature, the second hand feature and the third hand feature comprises:

concatenating the first hand feature, the second hand feature and the third hand feature to obtain the hand feature corresponding to the hand image.

6 . The method of claim 1 , wherein the interaction feature is determined by:

performing interaction on the respective hand features corresponding to the hand images based on a predetermined cross-view attention algorithm and/or a predetermined view-sharing algorithm, to obtain the interaction feature.

7 . The method of claim 6 , wherein the interaction feature is determined based on the cross-view attention algorithm by:

determining an attention score between hand key points corresponding to the hand images; and

performing the interaction on the respective hand features corresponding to the hand images by using the attention score to obtain the interaction feature.

8 . The method of claim 6 , wherein the interaction feature is determined based on the view-sharing algorithm by:

determining features of respective hand key points with the highest response at different angles of view using maximum value pooling; and

performing the interaction on the respective hand features corresponding to the hand images by using the features of the respective hand key points with the highest response at different angles of view to obtain the interaction feature.

9 . The method of claim 1 , further comprising:

determining a loss value using a predetermined loss function based on the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results; and

adjusting a network parameter of the hand three-dimensional reconstruction network by using the loss value, to obtain an adjusted hand three-dimensional reconstruction network.

10 . The method of claim 9 , wherein determining the loss value using the predetermined loss function based on the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results comprises:

determining, as the loss value, a difference between the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results.

11 . The method of claim 9 , wherein determining the loss value using the predetermined loss function based on the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results comprises:

for each angle of view of the at least two angles of view,

projecting an initial hand three-dimensional model corresponding to the angle of view onto the hand image of the angle of view to obtain a first projection point set;

rotating, to the angle of view, an initial hand three-dimensional model corresponding to an angle of view other than the angle of view;

projecting the rotated hand three-dimensional model onto the hand image at the angle of view to obtain a second projection point set; and

determining a difference between the first projection point set and the second projection point set as the loss value.

12 . The method of claim 9 , wherein determining the loss value using the predetermined loss function based on the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results comprises:

identifying, as a pseudo tag, hand key points in the hand images acquired at the at least two angles of view by using a predefined hand key point positioning algorithm; and

determining a difference between the pseudo tag and an initial hand key point as the loss value.

13 . The method of claim 9 , wherein determining the loss value using the predetermined loss function based on the fused hand three-dimensional reconstruction result and the initial hand three-dimensional reconstruction results comprises:

obtaining a predetermined hand three-dimensional model reference template; and

determining, as the loss value, a difference between the hand three-dimensional model reference template and an initial hand three-dimensional model.

14 . An electronic device, comprising: one or more processors; a tangible storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a method of hand three-dimensional reconstruction, comprising:

obtaining hand images acquired at at least two angles of view;

determining, by using an interaction feature, an initial hand three-dimensional reconstruction result corresponding to each of the hand images based on a predetermined hand three-dimensional reconstruction network, wherein the interaction feature is obtained by performing interaction on hand features corresponding to each of the hand images, and the initial hand three-dimensional reconstruction result comprises a hand three-dimensional model and a hand key point; and

fusing the initial hand three-dimensional reconstruction results corresponding to the hand images acquired at the at least two angles of view to obtain a fused hand three-dimensional reconstruction result.

15 . The device of claim 14 , wherein determining, by using the interaction feature, the initial hand three-dimensional reconstruction result corresponding to each of the hand images based on the predetermined hand three-dimensional reconstruction network comprises:

for each hand image of the hand images,

determining a hand feature corresponding to the hand image based on the predetermined hand three-dimensional reconstruction network;

updating the hand feature corresponding to the hand image by using the interaction feature to obtain an updated hand feature corresponding to the hand image; and

determining the initial hand three-dimensional reconstruction result corresponding to the hand image based on the hand three-dimensional reconstruction network and the updated hand feature corresponding to the hand image.

16 . The device of claim 15 , wherein determining the hand feature corresponding to the hand image based on the predetermined hand three-dimensional reconstruction network comprises:

determining target information by inputting the hand image into the predetermined hand three-dimensional reconstruction network, wherein the target information comprises a feature map of at least one level and the hand key point; and

determining the hand feature corresponding to the hand image based on the target information.

17 . The device of claim 16 , wherein the target information comprises a hand gesture vector; and

determining the hand feature corresponding to the hand image based on the target information comprising:

determining the hand feature corresponding to the hand image based on at least one of a first hand feature, a second hand feature and a third hand feature; and wherein

the first hand feature is obtained by encoding the hand gesture vector and a coordinate of the hand key point;

the second hand feature is determined by processing a feature map of a target level using a predefined graph algorithm; and

the third hand feature is obtained by projecting the hand key point onto feature maps of levels other than the feature map of the target level, and determining features at projection position points on respective feature maps.

18 . The device of claim 17 , wherein determining the hand feature corresponding to the hand image based on at least one of the first hand feature, the second hand feature and the third hand feature comprises:

concatenating the first hand feature, the second hand feature and the third hand feature to obtain the hand feature corresponding to the hand image.

19 . The device of claim 15 , wherein the interaction feature is determined by:

performing interaction on the respective hand features corresponding to the hand images based on a predetermined cross-view attention algorithm and/or a predetermined view-sharing algorithm, to obtain the interaction feature.

20 . A tangible computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor of a dedicated hardware-based system, a method of hand three-dimensional reconstruction, comprising:

obtaining hand images acquired at at least two angles of view;

determining, by using an interaction feature, an initial hand three-dimensional reconstruction result corresponding to each of the hand images based on a predetermined hand three-dimensional reconstruction network, wherein the interaction feature is obtained by performing interaction on hand features corresponding to each of the hand images, and the initial hand three-dimensional reconstruction result comprises a hand three-dimensional model and a hand key point; and

fusing the initial hand three-dimensional reconstruction results corresponding to the hand images acquired at the at least two angles of view to obtain a fused hand three-dimensional reconstruction result.