IP Library Granted Patent US 12,548,274
Granted Patent B2
US 12,548,274 · App. 18/581,670 · Granted Feb 10, 2026

Method, device, and computer program product for generating 3D object reconstruction model

Inventors: Zijia Wang (London, GB); Zhisong Liu (Shenzhen, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T19/20G06T7/40G06T7/50G06T7/73G06T17/00G06T2207/20081G06T2207/20084G06T2207/30201G06T2219/2024G06V10/774G06V10/776G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,274
App. No.
18/581,670
Granted
Feb 10, 2026
Kind
B2
Abstract

The present disclosure relates to a method, a device, and a computer program product for generating a three-dimensional (3D) object reconstruction model. The method includes acquiring an input image containing a two-dimensional (2D) object. The method further includes determining a shape feature, a texture feature, and a posture feature of the 2D object. The method further includes generating a rendered image based on the shape feature, the texture feature, the posture feature, and the input image. The method further includes generating a 3D object reconstruction model based on the input image and the rendered image, and tuning the 3D object reconstruction model according to a parameter tuning model for stylization. The 3D object reconstruction model generated by this method can realistically reconstruct the 2D object in terms of the shape, the texture, and the posture, so that the 3D object outputted from the model can better match the input image.

Claims (72)

1 . A method comprising:

acquiring an input image containing a two-dimensional (2D) object;

determining a shape feature, a texture feature, and a posture feature of the 2D object;

generating a rendered image based on the shape feature, the texture feature, the posture feature, and the input image;

generating a 3D object reconstruction model based on the input image and the rendered image, wherein the 3D object reconstruction model is configured to generate a 3D object from the 2D object utilizing a weighted combination of multiple distinct loss functions including at least one loss function that is generated based on (i) estimated expression coefficients of the 3D object reconstruction model and (ii) an output of an expression classifier of the 3D object reconstruction model; and

tuning the 3D object reconstruction model according to a parameter tuning model for stylization.

2 . The method according to claim 1 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining an average object shape feature, and an identity feature and an expression feature of the 2D object; and

determining the shape feature based on the average object shape feature, the identity feature, and the expression feature as determined.

3 . The method according to claim 1 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining an average object texture feature and a constant feature of the 2D object; and

determining the texture feature based on the average object texture feature and the constant feature of the 2D object as determined.

4 . The method according to claim 1 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining a rotation feature, a scaling feature, and a translation feature of the 2D object; and

determining the posture feature based on the rotation feature, the scaling feature, and the translation feature as determined.

5 . The method according to claim 1 , wherein generating the 3D object reconstruction model comprises:

determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility based on the shape feature, the texture feature, and the posture feature; and

generating the 3D object reconstruction model based on the reconstruction loss, the expression loss, and the regularization loss.

6 . The method according to claim 5 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

determining input features of the input image and rendered features of the rendered image; and

generating the reconstruction loss based on the input image, the rendered image, the input features, and the rendered features.

7 . The method according to claim 5 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

generating an emotional expression coefficient based on the shape feature, the texture feature, and the posture feature;

extracting a real expression label of the 2D object; and

generating the expression loss based on the emotional expression coefficient and the real expression label.

8 . The method according to claim 5 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

determining an average shape feature, an average texture feature, and an average posture feature based on average features of the object; and

generating the regularization loss based on the average shape feature, the average texture feature, the average posture feature, the shape feature, the texture feature, and the posture feature.

9 . The method according to claim 1 , wherein tuning the 3D object reconstruction model comprises:

taking the input image, the rendered image, and a style label of the parameter tuning model as observed variables; and

tuning the 3D object reconstruction model based on a lower bound of log marginal likelihood of the observed variables.

10 . The method according to claim 1 , further comprising:

reconstructing the 2D object into a 3D object based on the tuned 3D object reconstruction model.

11 . An electronic device, comprising:

at least one processor; and

memory coupled to the at least one processor and having instructions stored therein, wherein the instructions, when executed by the at least one processor, cause the electronic device to perform actions comprising:

acquiring an input image containing a two-dimensional (2D) object;

determining a shape feature, a texture feature, and a posture feature of the 2D object;

generating a rendered image based on the shape feature, the texture feature, the posture feature, and the input image;

generating a 3D object reconstruction model based on the input image and the rendered image, wherein the 3D object reconstruction model is configured to generate a 3D object from the 2D object utilizing a weighted combination of multiple distinct loss functions including at least one loss function that is generated based on (i) estimated expression coefficients of the 3D object reconstruction model and (ii) an output of an expression classifier of the 3D object reconstruction model; and

tuning the 3D object reconstruction model according to a parameter tuning model for stylization.

12 . The electronic device according to claim 11 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining an average object shape feature, and an identity feature and an expression feature of the 2D object; and

determining the shape feature based on the average object shape feature, the identity feature, and the expression feature as determined.

13 . The electronic device according to claim 11 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining an average object texture feature and a constant feature of the 2D object; and

determining the texture feature based on the average object texture feature and the constant feature of the 2D object as determined.

14 . The electronic device according to claim 11 , wherein determining a shape feature, a texture feature, and a posture feature of the 2D object comprises:

determining a rotation feature, a scaling feature, and a translation feature of the 2D object; and

determining the posture feature based on the rotation feature, the scaling feature, and the translation feature as determined.

15 . The electronic device according to claim 11 , wherein generating the 3D object reconstruction model comprises:

determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility based on the shape feature, the texture feature, and the posture feature; and

generating the 3D object reconstruction model based on the reconstruction loss, the expression loss, and the regularization loss.

16 . The electronic device according to claim 15 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

determining input features of the input image and rendered features of the rendered image; and

generating the reconstruction loss based on the input image, the rendered image, the input features, and the rendered features.

17 . The electronic device according to claim 15 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

generating an emotional expression coefficient based on the shape feature, the texture feature, and the posture feature;

extracting a real expression label of the 2D object; and

generating the expression loss based on the emotional expression coefficient and the real expression label.

18 . The electronic device according to claim 15 , wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object plausibility comprises:

determining an average shape feature, an average texture feature, and an average posture feature based on average features of the object; and

generating the regularization loss based on the average shape feature, the average texture feature, the average posture feature, the shape feature, the texture feature, and the posture feature.

19 . The electronic device according to claim 11 , wherein tuning the 3D object reconstruction model comprises:

taking the input image, the rendered image, and a style label of the parameter tuning model as observed variables; and

tuning the 3D object reconstruction model based on a lower bound of log marginal likelihood of the observed variables.

20 . A computer program product comprising a non-transitory computer-readable medium having machine-executable instructions stored therein, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

acquiring an input image containing a two-dimensional (2D) object;

determining a shape feature, a texture feature, and a posture feature of the 2D object;

generating a rendered image based on the shape feature, the texture feature, the posture feature, and the input image;

generating a 3D object reconstruction model based on the input image and the rendered image, wherein the 3D object reconstruction model is configured to generate a 3D object from the 2D object utilizing a weighted combination of multiple distinct loss functions including at least one loss function that is generated based on (i) estimated expression coefficients of the 3D object reconstruction model and (ii) an output of an expression classifier of the 3D object reconstruction model; and

tuning the 3D object reconstruction model according to a parameter tuning model for stylization.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: WANG, ZIJIA; LIU, ZHISONG; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 066498/0022 →
Priority Claims (1)
CN 202410114542.5 · Jan 26, 2024 · national
Continuity (1)
Related Publication 20250245946A1 · Jul 31, 2025
References Cited (29)
US 11308657B1 · Berlin · 2022 [cited by examiner]
US 11580395B2 · Karras · 2023 [cited by examiner]
US 11734890B2 · Laine · 2023 [cited by examiner]
US 11861475B2 · Hu · 2024 [cited by examiner]
US 12392665B1 · Guan · 2025 [cited by examiner]
US 12400631B2 · Wang · 2025 [cited by examiner]
US 20100214290A1 · Shiell · 2010 [cited by examiner]
US 20120301013A1 · Gu · 2012 [cited by examiner]
US 20210056348A1 · Berlin · 2021 [cited by examiner]
US 20220036635A1 · Li · 2022 [cited by examiner]
US 20220343522A1 · Bi · 2022 [cited by examiner]
US 20230005251A1 · Take · 2023 [cited by examiner]
US 20230243973A1 · Hung · 2023 [cited by examiner]
US 20240134937A1 · Ni · 2024 [cited by examiner]
US 20240169563A1 · Wen · 2024 [cited by examiner]
US 20240221316A1 · Gorodissky · 2024 [cited by examiner]
US 20240273871A1 · Song · 2024 [cited by examiner]
US 20240331322A1 · Smith · 2024 [cited by examiner]
US 20240404253A1 · Raguram · 2024 [cited by examiner]
US 20250245946A1 · Wang · 2025 [cited by examiner]
A. T. Tran et al., “Regressing Robust and Discriminative 3D Morphable Models with a Very Deep Neural Network,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Dec. 15, 2016, pp. 5163-5172. [cited by applicant]
K. Genova et al., “Unsupervised Training for 3D Morphable Model Regression,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 8377-8386. [cited by applicant]
L. Tran et al., “Towards High-Fidelity Nonlinear 3D Face Morphable Model,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Apr. 9, 2019, pp. 1126-1135. [cited by applicant]
V. Blanz et al., “A Morphable Model for the Synthesis of 3D Faces,” Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, Jul. 1999, 8 pages. [cited by applicant]
I. Kemelmacher-Shlizerman et al., “Face Reconstruction in the Wild,” International Conference on Computer Vision (ICCV), Nov. 2011, 8 pages. [cited by applicant]
J. Johnson et al., “Perceptual Losses for Real-time Style Transfer and Super-resolution,” European Conference on Computer Vision, Sep. 17, 2016, 17 pages. [cited by applicant]
K. Simonyan et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition,” International Conference on Learning Representations, arXiv:1409.1556v6, Apr. 10, 2015, 14 pages. [cited by applicant]
A. Mollahosseini et al., “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild,” arXiv:1708.03985v4, Oct. 9, 2017, 18 page. [cited by applicant]
T. Gerig et al., “Morphable Face Models—An Open Framework,” arXiv:1709.08398v2, Sep. 26, 2017, 8 pages. [cited by applicant]