IP Library › Granted Patent US 12,260,492
Granted Patent B2
US 12,260,492 · App. 18/099,602 · Granted Mar 25, 2025

Method and apparatus for training a three-dimensional face reconstruction model and method and apparatus for generating a three-dimensional face image

Inventors: Di Wang (Beijing, CN); Ruizhi Chen (Beijing, CN); Chen Zhao (Beijing, CN); Jingtuo Liu (Beijing, CN); Errui Ding (Beijing, CN); Tian Wu (Beijing, CN); Haifeng Wang (Beijing, CN)
Assignee: Beijing Baidu Netcom Science Technology Co., Ltd.
G06T15/20G06T15/04G06T15/40G06T19/20G06V10/774G06V10/82G06V40/168G06T2219/2004G06T2219/2012G06T2219/2016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,492
App. No.
18/099,602
Granted
Mar 25, 2025
Kind
B2
Abstract

A method for training a three-dimensional face reconstruction model includes inputting an acquired sample face image into a three-dimensional face reconstruction model to obtain a coordinate transformation parameter and a face parameter of the sample face image; determining the three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the acquired stylized face map of the sample face image; transforming the three-dimensional stylized face image of the sample face image into a camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain a rendered map; and training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image.

Claims (86)

1. A method for training a three-dimensional face reconstruction model, comprising:

acquiring a sample face image and a stylized face map of the sample face image;

inputting the sample face image into a three-dimensional face reconstruction model to obtain a coordinate transformation parameter and a face parameter of the sample face image;

determining a three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image;

transforming the three-dimensional stylized face image of the sample face image into a camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain a rendered map; and

training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image;

wherein determining the three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image comprises:

constructing a three-dimensional face image of the sample face image based on the face parameter of the sample face image; and

processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image;

wherein processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image comprises:

performing texture expansion on the stylized face map of the sample face image to obtain an initial texture map;

performing at least one of occlusion removal processing, highlight removal processing, or face pose adjustment processing on the initial texture map based on a map regression network to obtain a target texture map, wherein the map regression network is a pre-trained convolutional neural network for processing the initial texture map; and

processing the three-dimensional face image of the sample face image according to the target texture map to obtain the three-dimensional stylized face image of the sample face image.

2. The method of claim 1 , wherein acquiring the stylized face map of the sample face image comprises:

extracting a stylized feature from a stylized coding network;

inputting the sample face image into a face restoration coding network to obtain a face feature of the sample face image; and

generating, based on a style map generation network, the stylized face map of the sample face image according to the stylized feature and the face feature of the sample face image.

3. The method of claim 1 , wherein training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image comprises:

jointly training the three-dimensional face reconstruction model and the map regression network according to the rendered map and the stylized face map of the sample face image.

4. The method of claim 1 , wherein training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image comprises:

extracting a stylized face region from the stylized face map of the sample face image;

adjusting a background color of the stylized face region according to a background color of the rendered map;

determining an image comparison loss according to the rendered map and the adjusted stylized face region; and

training the three-dimensional face reconstruction model according to the image comparison loss.

5. The method of claim 1 , wherein inputting the sample face image into the three-dimensional face reconstruction model to obtain the coordinate transformation parameter and the face parameter of the sample face image comprises:

inputting the sample face image into the three-dimensional face reconstruction model to obtain Euler angles, a translation transformation parameter and a scaling transformation parameter in the coordinate transformation parameter, and the face parameter of the sample face image; and

correspondingly, transforming the three-dimensional stylized face image of the sample face image into the camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain the rendered map comprises:

affinely transforming the three-dimensional stylized face image of the sample face image into the camera coordinate system based on the translation transformation parameter and the scaling transformation parameter; and

rendering, based on a field of view of a camera and the Euler angles, the three-dimensional stylized face image subjected to the affine transformation to obtain the rendered map.

6. The method of claim 1 , wherein the face parameter of the sample face image comprises a face shape parameter.

7. A method for generating a three-dimensional face image, comprising:

acquiring a target face image and a stylized face map of the target face image;

inputting the target face image into a three-dimensional face reconstruction model to obtain a face parameter of the target face image, wherein the three-dimensional face reconstruction model is obtained through training based on the method for training a three-dimensional face reconstruction model of claim 1 ; and

determining a three-dimensional stylized face image of the target face image according to the face parameter of the target face image and the stylized face map of the target face image.

8. The method of claim 7 , wherein acquiring the stylized face map of the target face image comprises:

extracting a stylized feature from the stylized coding network;

inputting the target face image into a face restoration coding network to obtain a face feature of the target face image; and

generating, based on a style map generation network, the stylized face map of the target face image according to the stylized feature and the face feature of the target face image.

9. The method of claim 7 , wherein determining the three-dimensional stylized face image of the target face image according to the face parameter of the target face image and the stylized face map of the target face image comprises:

constructing a three-dimensional face image of the target face image based on the face parameter of the target face image; and

processing the three-dimensional face image of the target face image according to the stylized face map of the target face image to obtain the three-dimensional stylized face image of the target face image.

10. The method of claim 9 , wherein processing the three-dimensional face image of the target face image according to the stylized face map of the target face image to obtain the three-dimensional stylized face image of the target face image comprises:

performing texture expansion on the stylized face map of the target face image to obtain a to-be-processed texture map;

performing at least one of occlusion removal processing, highlight removal processing, or face pose adjustment processing on the to-be-processed texture map based on the map regression network to obtain a processed texture map; and

processing the three-dimensional face image of the target face image according to the processed texture map to obtain the three-dimensional stylized face image of the target face image.

11. The method of claim 7 , wherein the face parameter of the target face image comprises a face shape parameter.

12. A non-transitory computer-readable storage medium storing a computer instruction, wherein the computer instruction is configured to cause a computer to execute the method for training a three-dimensional face reconstruction model of claim 7 .

13. An electronic device, comprising:

at least one processor; and

a memory communicatively connected to the at least one processor,

wherein the memory stores an instruction executable by the at least one processor, and the instruction is executed by the at least one processor to cause the at least one processor to execute the following steps:

acquiring a sample face image and a stylized face map of the sample face image;

inputting the sample face image into a three-dimensional face reconstruction model to obtain a coordinate transformation parameter and a face parameter of the sample face image;

determining a three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image;

transforming the three-dimensional stylized face image of the sample face image into a camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain a rendered map; and

training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image;

wherein determining the three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image comprises:

constructing a three-dimensional face image of the sample face image based on the face parameter of the sample face image; and

processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image;

wherein processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image comprises:

performing texture expansion on the stylized face map of the sample face image to obtain an initial texture map;

performing at least one of occlusion removal processing, highlight removal processing, or face pose adjustment processing on the initial texture map based on a map regression network to obtain a target texture map, wherein the map regression network is a pre-trained convolutional neural network for processing the initial texture map; and

processing the three-dimensional face image of the sample face image according to the target texture map to obtain the three-dimensional stylized face image of the sample face image.

14. The electronic device of claim 13 , wherein acquiring the stylized face map of the sample face image comprises:

extracting a stylized feature from a stylized coding network;

inputting the sample face image into a face restoration coding network to obtain a face feature of the sample face image; and

generating, based on a style map generation network, the stylized face map of the sample face image according to the stylized feature and the face feature of the sample face image.

15. An electronic device, comprising:

at least one processor; and

a memory communicatively connected to the at least one processor, wherein the memory stores an instruction executable by the at least one processor, and the instruction is executed by the at least one processor to cause the at least one processor to execute the following steps:

acquiring a target face image and a stylized face map of the target face image;

inputting the target face image into a three-dimensional face reconstruction model to obtain a face parameter of the target face image, wherein the three-dimensional face reconstruction model is obtained through training based on the electronic device of claim 13 ; and

determining a three-dimensional stylized face image of the target face image according to the face parameter of the target face image and the stylized face map of the target face image.

16. A non-transitory computer-readable storage medium storing a computer instruction, wherein the computer instruction is configured to cause a computer to execute the following steps:

acquiring a sample face image and a stylized face map of the sample face image;

inputting the sample face image into a three-dimensional face reconstruction model to obtain a coordinate transformation parameter and a face parameter of the sample face image;

determining a three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image;

transforming the three-dimensional stylized face image of the sample face image into a camera coordinate system based on the coordinate transformation parameter, and rendering the transformed three-dimensional stylized face image to obtain a rendered map; and

training the three-dimensional face reconstruction model according to the rendered map and the stylized face map of the sample face image;

wherein determining the three-dimensional stylized face image of the sample face image according to the face parameter of the sample face image and the stylized face map of the sample face image comprises:

constructing a three-dimensional face image of the sample face image based on the face parameter of the sample face image; and

processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image;

wherein processing the three-dimensional face image of the sample face image according to the stylized face map of the sample face image to obtain the three-dimensional stylized face image of the sample face image comprises:

performing texture expansion on the stylized face map of the sample face image to obtain an initial texture map;

performing at least one of occlusion removal processing, highlight removal processing, or face pose adjustment processing on the initial texture map based on a map regression network to obtain a target texture map, wherein the map regression network is a pre-trained convolutional neural network for processing the initial texture map; and

processing the three-dimensional face image of the sample face image according to the target texture map to obtain the three-dimensional stylized face image of the sample face image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2023
From: WANG, DI; CHEN, RUIZHI; ZHAO, CHEN; LIU, JINGTUO; DING, ERRUI; WU, TIAN; WANG, HAIFENG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 062764/0177 →
Priority Claims (1)
CN 202210738050.4 · Jun 28, 2022 · national
Continuity (1)
Related Publication 20230419592A1 · Dec 28, 2023
References Cited (21)
US 10636175B2 · Caballero · 2020 [cited by examiner]
US 10650227B2 · Cole · 2020 [cited by examiner]
US 10991154B1 · Li et al. · 2021 [cited by applicant]
US 11443460B2 · Caballero · 2022 [cited by examiner]
US 11682098B2 · Yip · 2023 [cited by examiner]
US 11854203B1 · Gafni · 2023 [cited by examiner]
US 12014829B2 · Kramer · 2024 [cited by examiner]
US 20190035149A1 · Chen et al. · 2019 [cited by applicant]
US 20220383580A1 · Rohmetra · 2022 [cited by examiner]
US 20230377324A1 · Kim · 2023 [cited by examiner]
US 20240135627A1 · Song · 2024 [cited by examiner]
US 20240242408A1 · Gudkov · 2024 [cited by examiner]
CN 108399649A · 2018 [cited by applicant]
CN 109255830A · 2019 [cited by applicant]
CN 111951372A · 2020 [cited by applicant]
CN 113506367A · 2021 [cited by applicant]
CN 113870420A · 2021 [cited by applicant]
Shiri F, Yu X, Porikli F, Hartley R, Koniusz P. Identity-preserving face recovery from stylized portraits. International Journal of Computer Vision. Jun. 1, 2019;127:863-83. [cited by examiner]
Men Y, Liu H, Yao Y, Cui M, Xie X, Lian Z. 3DToonify: Creating Your High-Fidelity 3D Stylized Avatar Easily from 2D Portrait Images. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 20… [cited by examiner]
Han F, Ye S, He M, Chai M, Liao J. Exemplar-based 3d portrait stylization. IEEE Transactions on Visualization and Computer Graphics. Sep. 24, 2021;29(2):1371-83. [cited by examiner]
Notice of Reasons for Refusal dated Nov. 5, 2023 issued in Japanese patent application No. 2023-029589. [cited by applicant]