IP Library Granted Patent US 12,657,879
Granted Patent B2
US 12,657,879 · App. 18/168,867 · Granted Jun 16, 2026

Multi-dimensional image stylization using transfer learning

Inventors: Guoxian Song (Los Angeles, CA); Hongyi Xu (Los Angeles, CA); Jing Liu (Los Angeles, CA); Tiancheng Zhi (Los Angeles, CA); Yichun Shi (Los Angeles, CA); Jianfeng Zhang (Singapore, SG); Zihang Jiang (Singapore, SG); Jiashi Feng (Singapore, SG); Shen Sang (Los Angeles, CA); Linjie Luo (Los Angeles, CA)
G06V10/7715G06V10/28G06V10/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,879
App. No.
18/168,867
Granted
Jun 16, 2026
Kind
B2
Abstract

A method for generating a multi-dimensional stylized image. The method includes providing input data into a latent space for a style conditioned multi-dimensional generator of a multi-dimensional generative model and generating the multi-dimensional stylized image from the input data by the style conditioned multi-dimensional generator. The method further includes synthesizing content for the multi-dimensional stylized image using a latent code and corresponding camera pose from the latent space to formulate an intermediate code to modulate synthesis convolution layers to generate feature images as multi-planar representations and synthesizing stylized feature images of the feature images for generating the multi-dimensional stylized image of the input data. The style conditioned multi-dimensional generator is tuned using a guided transfer learning process using a style prior generator.

Claims (47)

1 . A method for generating a multi-dimensional stylized image, the method comprising:

providing input data into a latent space for a style conditioned multi-dimensional generator of a multi-dimensional generative model; and

generating the multi-dimensional stylized image from the input data by the style conditioned multi-dimensional generator by:

synthesizing content for the multi-dimensional stylized image using a latent code and corresponding camera pose from the latent space to formulate an intermediate code to modulate synthesis convolution layers to generate feature images as multi-planar representations; and

synthesizing stylized feature images of the feature images for generating the multi-dimensional stylized image of the input data,

wherein the style conditioned multi-dimensional generator is tuned using a guided transfer learning process using a style prior generator, wherein the style prior generator is a two-dimensional generator trained on style exemplars having a stylization and configured to generated stylized two-dimensional images of realistic images, and wherein the style prior generator is conditioned using a transfer learning process using a predetermined amount of style exemplars by synthesizing the style exemplars into feature maps having the stylization, wherein the predetermined amount of style exemplars is 10 or more.

2 . The method of claim 1 , wherein the guided transfer learning process includes:

generating two-dimensional stylized images of training data using the style prior generator;

generating the multi-dimensional stylized images of the training data using the style conditioned multi-dimensional generator; and

discriminating the two-dimensional stylized images and the multi-dimensional stylized images using a discriminator that compares a distribution of translated images of the multi-dimensional stylized images to a distribution of the two-dimensional stylized images to tune the style conditioned multi-dimensional generator to generate the stylized feature images.

3 . The method of claim 1 , wherein the feature images are further sampled into a neural radiance field at predetermined camera poses and accumulated to synthesize generation of raw feature images using volume rendering.

4 . The method of claim 3 , further comprising upsampling the volume rendered feature images, wherein the upsampling includes refining the volume rendered feature images using a super resolution module to synthesize RGB images of the multi-dimensional stylized image.

5 . The method of claim 1 , wherein the latent code and corresponding camera pose are passed through a mapping network to formulate the intermediate code.

6 . The method of claim 1 , wherein the predetermined amount of style exemplars is between 10 and 20.

7 . The method of claim 1 , wherein during the transfer learning process realistic images are used, wherein camera poses are determined from the realistic images.

8 . The method of claim 1 , wherein the stylization is an artistic style selected from the group consisting of cartoon, oil painting, cubism, abstract, comic, Sam Yang, sculpture, and anime.

9 . The method of claim 1 , wherein the providing the input data includes inverting the input data using an encoder that outputs latent codes, wherein the style conditioned multi-dimensional generator is configured to generate the multi-planar representations using the encoded intermediate latent code.

10 . The method of claim 9 , wherein the encoder is trained by:

randomly sampling a plurality of latent codes from a Gaussian-distributed latent space with a fixed front camera pose,

feeding the plurality of latent codes and the fixed front camera pose into a mapping network to obtain intermediate latent code samples,

training the encoder with the obtained intermediate latent code samples to encode intermediate latent codes.

11 . The method of claim 1 , wherein the multi-dimensional stylized image is a three-dimensional stylized portrait image and the input data is a self-portrait image.

12 . A computing system comprising:

one or more processors; and

one or more computer-readable media having computer-executable instructions stored thereon that, upon execution, cause the one or more processors to implement a multi-dimensional generative adversarial network (GAN), the GAN comprising:

an encoder for inverting input data into a latent space;

a style conditioned multi-dimensional generator; and

a style prior generator,

wherein the style conditioned multi-dimensional generator is configured to:

synthesize content for a multi-dimensional stylized image using a latent code and corresponding camera pose from the latent space to formulate an intermediate code to modulate synthesis convolution layers to generate feature images as multi-planar representations;

synthesize stylized feature images of the feature images as the multi-planar representations to generate the multi-dimensional stylized image of the input data,

wherein the style conditioned multi-dimensional generator is tuned using a guided transfer learning process using the style prior generator, wherein the style prior generator is a two-dimensional generator trained on style exemplars having a stylization and configured to generated stylized two-dimensional images of realistic images, and wherein the style prior generator is conditioned using a transfer learning process using a predetermined amount of style exemplars by synthesizing the style exemplars into feature maps having the stylization, wherein the predetermined amount of style exemplars is 10 or more.

13 . The computing system of claim 12 , further comprising a decoder, the decoder including a super resolution module to synthesize RGB images of the multi-dimensional stylized image.

14 . The computing system of claim 12 , further comprising a discriminator, wherein the style prior generator is configured to generate two-dimensional stylized images of training data and the discriminator is configured to discriminate between a distribution of the two-dimensional stylized images from the style prior generator and a distribution of translated images of multi-dimensional stylized images of the training data from the style conditioned multi-dimensional generator to tune the style conditioned multi-dimensional generator.

15 . The computing system of claim 14 , wherein the style prior generator is a two-dimensional generator trained on style exemplars having a stylization and configured to generate stylized two-dimensional images of realistic images.

16 . The computing system of claim 15 , wherein the stylization is an artistic style selected from the group consisting of cartoon, oil painting, cubism, abstract, comic, Sam Yang, sculpture, and anime.

17 . The computing system of claim 12 , wherein the multi-dimensional stylized image is a three-dimensional stylized portrait image and the input data is a self-portrait image.

18 . A non-transitory computer-readable medium having computer-executable instructions stored thereon that, upon execution, cause one or more processors to perform operations comprising:

providing input data into a latent space for a style conditioned multi-dimensional generator of a multi-dimensional generative model; and

generating a multi-dimensional stylized image from the input data by the style conditioned multi-dimensional generator by:

synthesizing content for the multi-dimensional stylized image using a latent code and corresponding camera pose from the latent space to formulate an intermediate code to modulate synthesis convolution layers to generate feature images as multi-planar representations;

synthesizing stylized feature images of the feature images for generating the multi-dimensional stylized image of the input data,

wherein the style conditioned multi-dimensional generator is tuned using a guided transfer learning process using a style prior generator, wherein the style prior generator is a two-dimensional generator trained on style exemplars having a stylization and configured to generated stylized two-dimensional images of realistic images, and wherein the style prior generator is conditioned using a transfer learning process using a predetermined amount of style exemplars by synthesizing the style exemplars into feature maps having the stylization, wherein the predetermined amount of style exemplars is 10 or more.

19 . The non-transitory computer-readable medium according to claim 18 , wherein the operations further include:

generating two-dimensional stylized images of training data using the style prior generator;

generating the multi-dimensional stylized images of the training data using the style conditioned multi-dimensional generator; and

discriminating the two-dimensional stylized images and the multi-dimensional stylized images using a discriminator that compares a distribution of translated images of the multi-dimensional stylized images to a distribution of the two-dimensional stylized images to tune the style conditioned multi-dimensional generator to generate the stylized feature images.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: ZHANG, JIANFENG; JIANG, ZIHANG; FENG, JIASHI
To: TIKTOK PTE. LTD.
Reel/Frame 074658/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: TIKTOK PTE. LTD.
To: LEMON INC.
Reel/Frame 074658/0771 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: XU, HONGYI; ZHI, TIANCHENG; SANG, SHEN; LUO, LINJIE
To: BYTEDANCE INC.
Reel/Frame 074660/0294 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 074660/0367 →
Continuity (1)
Related Publication 20240273871A1 · Aug 15, 2024
References Cited (20)
US 11170270B2 · Merler · 2021 [cited by examiner]
US 20210049468A1 · Karras · 2021 [cited by examiner]
US 20210150769A1 · Babaheidarian · 2021 [cited by examiner]
US 20210382936A1 · Tomar · 2021 [cited by examiner]
Chan Efficient Geometry-aware 3D Generative Adversarial Networks (Year: 2022). [cited by examiner]
Song AgileGAN: Stylizing Portraits by Inversion-Consistent Transfer Learning (Year: 2021). [cited by examiner]
Shwarz GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis (Year: 2021). [cited by examiner]
Chen et al.,“Efficient Geometry-aware 3D Generative Adversarial Networks”, Computer Science, Computer Vision and Pattern Recognition, arXiv:2112.07945, retrieved online from: https://arxiv.org/abs/2112.07945; Submitted … [cited by applicant]
Chong et al.,“JoJoGAN: One Shot Face Stylization”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://arxiv.org/abs/2112.11641; Submitted on Dec. 22, 2021 (v1), last revised Mar. 6… [cited by applicant]
Deng et al.,“ArcFace: Additive Angular Margin Loss for Deep Face Recognition”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://arxiv.org/abs/1801.07698; Submitted on Jan. 23, 20… [cited by applicant]
Gu et al.,“StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://arxiv.org/abs/2110.08985; Submitted … [cited by applicant]
Huang et al.,“Unsupervised Image-to-Image Translation via Pre-trained StyleGAN2 Network”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://www.semanticscholar.org/reader/084368bc… [cited by applicant]
Karras et al.,“Analyzing and Improving the Image Quality of StyleGAN”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from; https://arxiv.org/abs/1912.04958; Submitted on Dec. 3, 2019 (v1), … [cited by applicant]
Or-El et al.,“StyleSDF: High-Resolution 3D-Consistent Image and Geometry Generation”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://arxiv.org/abs/2112.11427; Submitted on Dec.… [cited by applicant]
Pinkney et al.,“Resolution Dependent GAN Interpolation for Controllable Image Synthesis Between Domains”, Computer Science, Computer Vision and Pattern Recognition, retrieved online from: https://arxiv.org/abs/2010.0533… [cited by applicant]
Song et al.,“AgileGAN: stylizing portraits by inversion-consistent transfer learning”, ACM Journals, ACM Transactions on Graphics, vol. 37, No. 4, Article 111, accessed online from: https://guoxiansong.github.io/homepag… [cited by applicant]
Tov et al.,“Designing an Encoder for StyleGAN Image Manipulation”, Computer Science, Computer Vision and Pattern Recognition, accessed online from: https://arxiv.org/abs/2102.02766; Submitted on Feb. 4, 2021. [cited by applicant]
Xu et al.,“Deep 3D Portrait from a Single Image”, Computer Science, Computer Vision and Pattern Recognition, accessed online from: https://arxiv.org/pdf/2004.11598.pdf; Submitted on Apr. 24, 2020. [cited by applicant]
Yang et al.,“Pastiche Master: Exemplar-Based High-Resolution Portrait Style Transfer”, Computer Science, Computer Vision and Pattern Recognition, accessed online from: https://arxiv.org/abs/2203.13248; Submitted on Mar.… [cited by applicant]
Zhang et al.,“The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, Computer Science, Computer Vision and Pattern Recognition, accessed online from: https://arxiv.org/abs/1801.03924; Submitted on Jan.… [cited by applicant]