IP Library › Granted Patent US 11,720,994
Granted Patent B2
US 11,720,994 · App. 17/321,384 · Granted Aug 8, 2023

High-resolution portrait stylization frameworks using a hierarchical variational encoder

Inventors: Linjie Luo (Los Angeles, CA); Guoxian Song (Singapore, SG); Jing Liu (Los Angeles, CA); Wanchun Ma (Los Angeles, CA)
Assignee: Lemon Inc.
G06T3/0012G06F18/214G06N3/045G06N3/08G06T3/0006G06T5/00G06T11/00G06T2207/20016G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,994
App. No.
17/321,384
Granted
Aug 8, 2023
Kind
B2
Abstract

Systems and method directed to an inversion-consistent transfer learning framework for generating portrait stylization using only limited exemplars. In examples, an input image is received and encoded using a variational autoencoder to generate a latent vector. The latent vector may be provided to a generative adversarial network (GAN) generator to generate a stylized image. In examples, the variational autoencoder is trained using a plurality of images while keeping the weights of a pre-trained GAN generator fixed, where the pre-trained GAN generator acts as a decoder for the encoder. In other examples, a multi-path attribute aware generator is trained using a plurality of exemplar images and learning transfer using the pre-trained GAN generator.

Claims (65)

1. A method for generating a stylized image, the method comprising:

receiving an input image;

encoding the input image using a variational autoencoder to obtain a latent vector by:

passing the received input image through a headless pyramid network to produce multiple levels of features maps at different sizes;

encoding, for each of the levels of features maps at different sizes, each level's respective feature map at the different size with a separate encoder of a plurality of encoders to produce a code, and

combining the encoded code of each level's respective feature map to obtain the latent vector;

providing the latent vector to a pre-trained generative adversarial network (GAN) model;

generating, by the pre-trained GAN model, a stylized image from the pre-trained GAN model, the generated stylized image being a cartoon style image of the input image; and

providing the stylized image as an output,

wherein the pre-trained GAN model includes a multi-path structure corresponding to two or more different attributes.

2. The method of claim 1 , further comprising:

receiving a plurality of exemplar images;

training a GAN model using transfer learning based on the received plurality of exemplar images; and

terminating the process of training when the output of the GAN model satisfies a predetermined condition at a first time to produce the pre-trained GAN model.

3. The method of claim 2 , further comprising:

receiving a plurality of training images; and

training the variational autoencoder while keeping the weights of the pre-trained GAN model fixed.

4. The method of claim 1 , wherein the latent vector is sampled from a standard Gaussian distribution.

5. The method of claim 4 , further comprising:

mapping the latent vector to an intermediate vector; and

forwarding the intermediate vector to an affine transform within a style block of the pre-trained GAN model.

6. The method of claim 1 , wherein the pre-trained GAN model comprises a pre-trained StyleGAN2 model.

7. A system configured to generate a stylized image, the system comprising:

a processor; and

memory including instructions, which when executed by the processor, causes the processor to:

receive an input image;

encode the input image using a variational autoencoder to obtain a latent vector by:

passing the received input image through a headless pyramid network to produce multiple levels of features maps at different sizes;

encoding, for each of the levels of features maps at different sizes, each level's respective feature map at the different size with a separate encoder of a plurality of encoders to produce a code, and

combining the encoded code of each level's respective feature map to obtain the latent vector;

provide the latent vector to a pre-trained generative adversarial network (GAN) model;

generate, by the pre-trained GAN model, a stylized image from the pre-trained GAN model, the generated stylized image being a cartoon style image of the input image; and

provide the stylized image as an output,

wherein the pre-trained GAN model includes a multi-path structure corresponding to two or more different attributes.

8. The system of claim 7 , wherein the instructions, when executed by the processor, cause the processor to:

receive a plurality of exemplar images;

train the GAN model using transfer learning based on a pre-trained GAN model and the received plurality of exemplar images and

terminate the process of training when the output of the GAN model satisfies a predetermined condition at a first time to produce the pre-trained GAN model.

9. The system of claim 8 , wherein the instructions, when executed by the processor, cause the processor to:

receive a plurality of training images; and

training the variational autoencoder while keeping the weights of the pre-trained GAN model fixed.

10. The system of claim 7 , wherein the latent vector is sampled from a standard Gaussian distribution.

11. The system of claim 10 , wherein the instructions, when executed by the processor, cause the processor to:

map the latent vector to an intermediate vector; and

forward the intermediate vector to an affine transform within a style block of the pre-trained GAN model.

12. The system of claim 7 , wherein the pre-trained GAN model comprises a pre-trained StyleGAN2 model.

13. A non-transitory computer-readable storage medium including instructions, which when executed by a processor, cause the processor to:

receive an input image;

encode the input image using a variational autoencoder to obtain a latent vector by:

passing the received input image through a headless pyramid network to produce multiple levels of features maps at different sizes;

encoding, for each of the levels of features maps at different sizes, each level's respective feature map at the different size with a separate encoder of a plurality of encoders to produce a code, and

combining the encoded code of each level's respective feature map to obtain the latent vector;

provide the latent vector to a pre-trained generative adversarial network (GAN) model;

generate, by the pre-trained GAN model, a stylized image from the pre-trained GAN model, the generated stylized image being a cartoon style image of the input image; and

provide the stylized image as an output,

wherein the pre-trained GAN model includes a multi-path structure corresponding to two or more different attributes.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, which when executed by a processor, cause the processor to:

map a latent vector sampled from a standard Gaussian distribution to an intermediate vector; and

forward the intermediate vector to an affine transform within a style block of the pre-trained GAN model.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the combined code from each level's respective feature map to obtain the latent vector is passed to fully connected layers to generate means and standard deviations representing Gaussian importance distribution in a Z+ space.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, which when executed by a processor, cause the processor to:

receive a plurality of exemplar images including cartoon characters;

train GAN model using transfer learning based on the received plurality of exemplar images; and

terminating the process of training after at most 1200 interactions to produce the pre-trained GAN model.

17. The non-transitory computer-readable storage medium of claim 13 , wherein the pre-trained GAN model comprises a pre-trained StyleGAN2 model.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: LUO, LINJIE; LIU, JING; MA, WANCHUN
To: BYTEDANCE INC.
Reel/Frame 056286/0766 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: TIKTOK PTE. LTD.
To: LEMON INC.
Reel/Frame 056286/0920 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 056286/0944 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: SONG, GUOXIAN
To: TIKTOK PTE. LTD.
Reel/Frame 057438/0321 →
Continuity (1)
Related Publication 20220375024A1 · Nov 24, 2022
Cited By (3)
US 12,272,031 US 12,652,404 US 12,657,833