IP Library › Granted Patent US 12,633,163
Granted Patent B2
US 12,633,163 · App. 18/042,149 · Granted May 19, 2026

Expression transformation method and apparatus, electronic device, and computer readable medium

Inventor: Qian He (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G06V40/174G06T11/00G06V10/7747G06V10/778
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,163
App. No.
18/042,149
Granted
May 19, 2026
Kind
B2
Abstract

An expression transformation method and apparatus, an electronic device, and a computer readable medium. The method comprises: acquiring a target face image ( 201 ); and inputting the target face image into a pre-trained expression transformation model to obtain an expression transformation image ( 202 ). The expression transformation model performs expression transformation on the target face image to achieve different expression transformation effects. A set of face images which are locally processed and have preset expressions displayed are used for training, so that the effect of the additional special effect can be achieved for the expression transformation image on the basis of transformation.

Claims (53)

1 . An expression transformation method, comprising:

obtaining a target face image;

inputting the target face image into an expression transformation model,

wherein the expression transformation model is pre-trained using an original face image set and an image set subjected to local processing and displaying preset expressions, and

wherein the image set subjected to local processing and displaying preset expressions for training the expression transformation model is obtained by performing local processing on particular regions of face images that present the preset expressions, and wherein the local processing performed on the particular regions of face images comprises image super-resolution processing and local whitening super-resolution processing; and

generating an expression transformation image by the expression transformation model performing expression transformation on the target face image, wherein the expression transformation image comprises a particular expression with an added effect, and wherein the added effect comprises beaming, whitened teeth, or colored hair.

2 . The method according to claim 1 , wherein after inputting the target face image into the pre-trained expression transformation model to obtain the expression transformation image, the method further comprises:

performing a masking processing on a target region of the expression transformation image to obtain an expression transformation image after the masking processing.

3 . The method according to claim 1 , wherein the expression transformation model comprises at least one of an expression transformation network or an expression transformation grid.

4 . The method according to claim 3 , further comprising:

obtaining the expression transformation image in response to inputting the target face image into the expression transformation network, wherein the expression transformation network is obtained through training by using the original face image set and the image set subjected to local processing and displaying the preset expressions.

5 . The method according to claim 4 , wherein the expression transformation network is obtained through training through the following steps:

obtaining the original face image set; and

inputting the original face image set and the image set subjected to local processing and displaying preset expressions into a preset first generative adversarial network for training to generate the expression transformation network.

6 . The method according to claim 1 , wherein the face images that present the preset expressions are obtained through:

inputting original face images in the original face image set into a pre-trained second generative adversarial network to obtain the face images that present the preset expressions.

7 . The method according to claim 3 , further comprising:

obtaining the expression transformation image corresponding to the target face image in response to inputting the target face image into the expression transformation grid, wherein the expression transformation grid is obtained through the original face image set and an original face image set subjected to local processing and displaying preset expressions.

8 . The method according to claim 3 , wherein the expression transformation grid is obtained through the following steps:

obtaining the original face image set;

performing local processing on each original face image that presents the preset expressions to obtain the image set subjected to local processing and displaying the preset expressions; and

storing, in preset grids, the original face image set and the face image set subjected to local processing and displaying the preset expressions to generate the expression transformation grid, wherein the expression transformation grid can represent a one-to-one corresponding relation between original face images and the expression transformation images corresponding thereto.

9 . An electronic device, comprising:

one or more processors; and

a storage apparatus, storing one or more programs, wherein

the one or more programs, when executed by the one or more processors, cause the one or more processors to implement operations comprising:

obtaining a target face image;

inputting the target face image into an expression transformation model,

wherein the expression transformation model is pre-trained using an original face image set and an image set subjected to local processing and displaying preset expressions, and

wherein the image set subjected to local processing and displaying preset expressions for training the expression transformation model is obtained by performing local processing on particular regions of face images that present the preset expressions, and wherein the local processing performed on the particular regions of face images comprises image super-resolution processing and local whitening super-resolution processing; and

generating an expression transformation image by the expression transformation model performing expression transformation on the target face image, wherein the expression transformation image comprises a particular expression with an added effect, and wherein the added effect comprises beaming, whitened teeth, or colored hair.

10 . A non-transitory computer readable medium, storing a computer program, wherein the program, when executed by a processor, implements operations comprising:

obtaining a target face image;

inputting the target face image into an expression transformation model,

wherein the expression transformation model is pre-trained using an original face image set and an image set subjected to local processing and displaying preset expressions, and

wherein the image set subjected to local processing and displaying preset expressions for training the expression transformation model is obtained by performing local processing on particular regions of face images that present the preset expressions, and wherein the local processing performed on the particular regions of face images comprises image super-resolution processing and local whitening super-resolution processing; and

generating an expression transformation image by the expression transformation model performing expression transformation on the target face image, wherein the expression transformation image comprises a particular expression with an added effect, and wherein the added effect comprises beaming, whitened teeth, or colored hair.

11 . The electronic device according to claim 9 , wherein after inputting the target face image into the pre-trained expression transformation model to obtain the expression transformation image, the operations further comprise:

performing a masking processing on a target region of the expression transformation image to obtain an expression transformation image after the masking processing.

12 . The electronic device according to claim 9 , wherein the expression transformation model comprises at least one of an expression transformation network or an expression transformation grid.

13 . The electronic device according to claim 12 , the operations further comprising:

obtaining the expression transformation image in response to inputting the target face image into the expression transformation network, wherein the expression transformation network is obtained through training by using the original face image set and the image set subjected to local processing and displaying the preset expressions.

14 . The electronic device according to claim 13 , wherein the expression transformation network is obtained through training through operations of:

obtaining the original face image set; and

inputting the original face image set and the image set subjected to local processing and displaying preset expressions into a preset first generative adversarial network for training to generate the expression transformation network.

15 . The electronic device according to claim 9 , wherein the face images that present the preset expressions are obtained through:

inputting original face images in the original face image set into a pre-trained second generative adversarial network to obtain the face images that present the preset expressions.

16 . The electronic device according to claim 12 , the operations further comprising:

obtaining the expression transformation image corresponding to the target face image in response to inputting the target face image into the expression transformation grid, wherein the expression transformation grid is obtained through the original face image set and the face image set subjected to local processing and displaying preset expressions.

17 . The electronic device according to claim 12 , wherein the expression transformation grid is obtained through operations of:

obtaining the original face image set;

performing local processing on each original face image that presents the preset expressions to obtain the image set subjected to local processing and displaying the preset expressions; and

storing, in preset grids, the original face image set and the face image set subjected to local processing and displaying the preset expressions to generate the expression transformation grid, wherein the expression transformation grid can represent a one-to-one corresponding relation between original face images and the expression transformation images corresponding thereto.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2023
From: HE, QIAN
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 062735/0820 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2023
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 062735/0982 →
Priority Claims (1)
CN 202010837052.X · Aug 19, 2020 · national
Continuity (1)
Related Publication 20230326248A1 · Oct 12, 2023
References Cited (25)
US 20070189627A1 · Cohen · 2007 [cited by examiner]
US 20130044958A1 · Brandt · 2013 [cited by examiner]
US 20220301348A1 · Bradley · 2022 [cited by examiner]
CN 107368810A · 2017 [cited by applicant]
CN 108958610A · 2018 [cited by applicant]
CN 109064393A · 2018 [cited by applicant]
CN 109859295A · 2019 [cited by applicant]
CN 110349232A · 2019 [cited by examiner]
CN 110490164A · 2019 [cited by applicant]
CN 110503601A · 2019 [cited by applicant]
CN 110689480A · 2020 [cited by examiner]
CN 111028305A · 2020 [cited by applicant]
CN 111179156A · 2020 [cited by applicant]
CN 111191564A · 2020 [cited by applicant]
CN 111243066A · 2020 [cited by applicant]
CN 111428572A · 2020 [cited by applicant]
CN 111968029A · 2020 [cited by applicant]
CN 117094893A · 2023 [cited by examiner]
Jiang, J., Ma, J., Chen, C., Jiang, X., & Wang, Z. (2017). Noise robust face image super-resolution through smooth sparse representation. IEEE Transactions on Cybernetics, 47(11), 3991-4002. https://doi.org/10.1109/tcyb… [cited by examiner]
Vahadane, A., Kumar, N., & Sethi, A. (2016). Learning based super-resolution of histological images. 2016 IEEE 13th International Symposium on Biomedical Imaging (ISBI), 816-819. https://doi.org/10.1109/isbi.2016.749339… [cited by examiner]
Written Opinion for International Application No. PCT/CN2021/113208, mailed Nov. 17, 2021, 8 Pages. [cited by applicant]
International Patent Application No. PCT/CN2021/113208; Int'l Search Report; dated Nov. 17, 2021; 2 pages. [cited by applicant]
European Patent Application No. 21857694.0; Extended Search Report; dated Jan. 3, 2024; 9 pages. [cited by applicant]
Choi et al.; “StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation”; IEEE/CVF Conf. on Computer Vision and Pattern Recognition; 2018; p. 8789-8797. [cited by applicant]
Lee et al.; “Facial Expression Transformations for Expression-Invariant Face Recognition”; Advances in Visual Computing Lecture Notes in Computer Science; 2006; p. 323-333. [cited by applicant]