IP Library › Granted Patent US 12,299,799
Granted Patent B2
US 12,299,799 · App. 18/046,073 · Granted May 13, 2025

Cascaded domain bridging for image generation

Inventors: Shen Sang (Los Angeles, CA); Tiancheng Zhi (Los Angeles, CA); Guoxian Song (Los Angeles, CA); Jing Liu (Los Angeles, CA); Linjie Luo (Los Angeles, CA); Chunpong Lai (Los Angeles, CA); Weihong Zeng (Beijing, CN); Jingna Sun (Beijing, CN); Xu Wang (Beijing, CN)
Assignees: Lemon Inc.; Beijing Zitiao Network Technology Co., Ltd.
G06T15/00G06T7/62G06V10/56G06V10/751G06V10/761G06T2207/10024G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,799
App. No.
18/046,073
Granted
May 13, 2025
Kind
B2
Abstract

A method of generating a stylized 3D avatar is provided. The method includes receiving an input image of a user, generating, using a generative adversarial network (GAN) generator, a stylized image, based on the input image, and providing the stylized image to a first model to generate a first plurality of parameters. The first plurality of parameters include a discrete parameter and a continuous parameter. The method further includes providing the stylized image and the first plurality of parameters to a second model that is trained to generate an avatar image, receiving, from the second model, the avatar image, comparing the stylized image to the avatar image, based on a loss function, to determine an error, updating the first model to generate a second plurality of parameters that correspond to the first plurality of parameters, based on the error, and providing the second plurality of parameters as an output.

Claims (68)

1. A method of generating a stylized 3D avatar, the method comprising:

receiving an input image of a user;

generating, using a generative adversarial network (GAN) generator, a stylized image, based on the input image;

providing the stylized image to a first model to generate a first plurality of parameters, the first plurality of parameters comprising one or more discrete parameters and one or more continuous parameters, the one or more continuous parameters comprise head characteristics, mouth characteristics, nose characteristics, ear characteristics, and eye characteristics, each corresponding to the user, wherein:

the head characteristics comprise a head width, a head length, and a blend shape coefficient for a head shape,

the mouth characteristics comprise a mouth width, a mouth volume, and a mouth position,

the nose characteristics comprise a nose width, a nose height, and a nose position,

the eye characteristics comprise an eye size, an eye spacing, and an eye rotation, and

the ear characteristics comprise an ear size,

providing the stylized image and the first plurality of parameters to a second model, the second model being trained to generate an avatar image;

receiving, from the second model, the avatar image;

comparing the stylized image to the avatar image, based on a loss function, to determine an error;

updating the first model to generate a second plurality of parameters, based on the error, the second plurality of parameters corresponding to the first plurality of parameters; and

providing the second plurality of parameters as an output.

2. The method of claim 1 , wherein the avatar image is a first image, wherein the error is a first error, and wherein the method further comprises:

receiving a predetermined loss threshold corresponding to the loss function;

providing the stylized image and the second plurality of parameters to the second model to generate a second avatar image;

comparing the stylized image to the second avatar image, based on the loss function, to determine a second error; and

providing the second plurality of parameters as the output, in response to determining that the second error is less than the predetermined loss threshold.

3. The method of claim 1 , wherein the loss function is based on a color loss corresponding to a difference between the stylized image and the avatar image, with respect to the one or more discrete parameters.

4. The method of claim 3 , wherein the loss function is further based on an identity loss and a perception loss, wherein the identity loss corresponds to a difference in a global appearance between the avatar and stylized images, wherein the avatar image and the stylized image each comprise a respective plurality of pixels, and wherein the perception loss corresponds to a difference between each pixel of the plurality of pixels of the stylized image to each respective pixel of the plurality of pixels of the avatar image.

5. The method of claim 1 , wherein the one or more discrete parameters comprise one or more from the group of: hair type, brow type, beard type, glasses type, eyelash type, eye makeup type, eye color type, brow color type, skin tone, hair color type, beard color type, mouth color type, and glasses color type.

6. A system for generating a stylized 3D avatar, the system comprising:

a processor; and

memory storing instructions that, when executed by the processor, cause the system to perform a set of operations, the set of operations comprising:

receiving an input image of a user;

generating, using a generative adversarial network (GAN) generator, a stylized image, based on the input image;

providing the stylized image to a first model to generate a first plurality of parameters, the first plurality of parameters comprising one or more discrete parameters and one or more continuous parameters, the one or more continuous parameters comprise head characteristics, mouth characteristics, nose characteristics, ear characteristics, and eye characteristics, each corresponding to the user, wherein:

the head characteristics comprise a head width, a head length, and a blend shape coefficient for a head shape,

the mouth characteristics comprise a mouth width, a mouth volume, and a mouth position,

the nose characteristics comprise a nose width, a nose height, and a nose position,

the eye characteristics comprise an eye size, an eye spacing, and an eye rotation, and

the ear characteristics comprise an ear size,

providing the stylized image and the first plurality of parameters to a second model, the second model being trained to generate an avatar image;

receiving, from the second model, the avatar image;

comparing the stylized image to the avatar image, based on a loss function, to determine an error;

updating the first model to generate a second plurality of parameters, based on the error, the second plurality of parameters corresponding to the first plurality of parameters; and

providing the second plurality of parameters as an output.

7. The system of claim 6 , wherein the avatar image is a first image, wherein the error is a first error, and wherein the set of operations further comprise:

receiving a predetermined loss threshold corresponding to the loss function;

providing the stylized image and the second plurality of parameters to the second model to generate a second avatar image;

comparing the stylized image to the second avatar image, based on the loss function, to determine a second error; and

providing the second plurality of parameters as the output, in response to determining that the second error is less than the predetermined loss threshold.

8. The system of claim 6 , wherein the loss function is based on a color loss corresponding to a difference between the stylized image and the avatar image, with respect to the one or more discrete parameters.

9. The system of claim 8 , wherein the loss function is further based on an identity loss and a perception loss, wherein the identity loss corresponds to a difference in a global appearance between the avatar and stylized images, wherein the avatar image and the stylized image each comprise a respective plurality of pixels, and wherein the perception loss corresponds to a difference between each pixel of the plurality of pixels of the stylized image to each respective pixel of the plurality of pixels of the avatar image.

10. The system of claim 6 , wherein the one or more discrete parameters comprise one or more from the group of: hair type, brow type, beard type, glasses type, eyelash type, eye makeup type, eye color type, brow color type, skin tone, hair color type, beard color type, mouth color type, and glasses color type.

11. A non-transient computer-readable storage medium comprising instructions being executable by one or more processors to cause the one or more processors to:

receive an input image of a user;

generate, using a generative adversarial network (GAN) generator, a stylized image, based on the input image;

provide the stylized image to a first model to generate a first plurality of parameters, the first plurality of parameters comprising one or more discrete parameters and one or more continuous parameters, the one or more continuous parameters comprise head characteristics, mouth characteristics, nose characteristics, ear characteristics, and eye characteristics, each corresponding to the user, wherein:

the head characteristics comprise a head width, a head length, and a blend shape coefficient for a head shape,

the mouth characteristics comprise a mouth width, a mouth volume, and a mouth position,

the nose characteristics comprise a nose width, a nose height, and a nose position,

the eye characteristics comprise an eye size, an eye spacing, and an eye rotation, and

the ear characteristics comprise an ear size,

provide the stylized image and the first plurality of parameters to a second model, the second model being trained to generate an avatar image;

receive, from the second model, the avatar image;

compare the stylized image to the avatar image, based on a loss function, to determine an error;

update the first model to generate a second plurality of parameters, based on the error, the second plurality of parameters corresponding to the first plurality of parameters; and

provide the second plurality of parameters as an output.

12. The non-transient computer-readable storage medium of claim 11 , wherein the avatar image is a first image, wherein the error is a first error, and wherein the set of instructions further cause the one or more processors to:

receive a predetermined loss threshold corresponding to the loss function;

provide the stylized image and the second plurality of parameters to the second model to generate a second avatar image;

compare the stylized image to the second avatar image, based on the loss function, to determine a second error; and

provide the second plurality of parameters as the output, in response to determining that the second error is less than the predetermined loss threshold.

13. The non-transient computer-readable storage medium of claim 11 , wherein the loss function is based on a color loss corresponding to a difference between the stylized image and the avatar image, with respect to the one or more discrete parameters.

14. The non-transient computer-readable storage medium of claim 13 , wherein the loss function is further based on an identity loss and a perception loss, wherein the identity loss corresponds to a difference in a global appearance between the avatar and stylized images, wherein the avatar image and the stylized image each comprise a respective plurality of pixels, and wherein the perception loss corresponds to a difference between each pixel of the plurality of pixels of the stylized image to each respective pixel of the plurality of pixels of the avatar image.

15. The non-transient computer-readable storage medium of claim 11 , wherein the one or more discrete parameters comprise one or more from the group of: hair type, brow type, beard type, glasses type, eyelash type, eye makeup type, eye color type, brow color type, skin tone, hair color type, beard color type, mouth color type, and glasses color type.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2025
From: SANG, SHEN; ZHI, TIANCHENG; SONG, GUOXIAN; LIU, JING; LUO, LINJIE; LAI, CHUNPONG
To: BYTEDANCE INC.
Reel/Frame 070838/0859 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2025
From: ZENG, WEIHONG
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 070839/0142 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2025
From: SUN, JINGNA; WANG, XU
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 070839/0333 →
Continuity (1)
Related Publication 20240135621A1 · Apr 25, 2024
References Cited (8)
US 20190295302A1 · Fu · 2019 [cited by examiner]
US 20210280322A1 · Frank · 2021 [cited by examiner]
US 20210358164A1 · Liu · 2021 [cited by applicant]
US 20210406682A1 · Meeyakhan · 2021 [cited by applicant]
US 20230130535A1 · Ma · 2023 [cited by examiner]
US 20230316606A1 · Qu · 2023 [cited by applicant]
US 20230342487A1 · Joseph · 2023 [cited by examiner]
Office Action dated Jun. 6, 2024 in U.S. Appl. No. 18/046,077 (26 pages). [cited by applicant]