IP Library › Granted Patent US 12,657,833
Granted Patent B2
US 12,657,833 · App. 18/832,859 · Granted Jun 16, 2026

Stylized image generation method and apparatus, electronic device and storage medium

Inventors: Wenyue Li (Beijing, CN); Caijin Zhou (Beijing, CN)
Assignee: Beijing Zitiao Network Technology Co., Ltd.
G06T19/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,833
App. No.
18/832,859
Granted
Jun 16, 2026
Kind
B2
Abstract

Embodiments of the present disclosure provide a stylized image generation method and apparatus, an electronic device, and a storage medium. The method comprises: determining a plurality of initial pairing data, and performing training on the basis of the initial pairing data to obtain a style model to be used; determining, on the basis of a preset screening condition, original images to be processed from original images, and processing, on the basis of said style model, each original image to be processed to obtain a style image to be used; performing deformation processing on the style image to be used to obtain a target style image, and taking each original image to be processed and the target style image corresponding thereto as stylized pairing data; and training, on the basis of the stylized pairing data, a stylized conversion model to be trained to obtain a target stylized conversion model.

Claims (71)

1 . A stylized image generation method, comprising:

determining a plurality of initial pairing data, and performing training on the basis of the plurality of initial pairing data to obtain a style model to be used, wherein each initial pairing data comprises an original image and an initial style image obtained after the original image is processed by a three dimensional (3D) style generation model;

determining a plurality of original images to be processed from the original images in the plurality of initial pairing data on the basis of a preset screening condition, and processing each original image to be processed on the basis of the style model to be used, so as to obtain a style image to be used corresponding to each original image to be processed;

obtaining a target style image corresponding to each original image to be processed by performing deformation processing on the style image to be used, and using each original image to be processed and the corresponding target style image as stylized pairing data; and

training a stylization conversion model to be trained on the basis of the stylized pairing data to obtain a target stylization conversion model, and upon acquiring video frames to be processed, performing stylization processing on the video frames to be processed on the basis of the target stylization conversion model to obtain a processed target video.

2 . The method according to claim 1 , wherein determining the plurality of initial pairing data comprises:

acquiring a plurality of original images comprising facial information; and

inputting each original image into a pre-trained 3D style generation model to obtain a corresponding initial style image after facial information processing.

3 . The method according to claim 1 , wherein performing training on the basis of the plurality of initial pairing data to obtain the style model to be used, comprises:

acquiring a first style model to be trained;

for each initial pairing data, using an original image in each initial pairing data as an input of the first style model to be trained to obtain a first output image corresponding to the original image;

determining a loss value on the basis of the first output image and the initial style image corresponding to the original image, so as to adjust model parameters in the first style model to be trained on the basis of the loss value; and

converging a first loss function in the first style model to be trained as a training target to obtain the style model to be used.

4 . The method according to claim 1 , wherein the preset screening condition comprises that change angles of parts to be adjusted are greater than preset change angle threshold values, and wherein determining the plurality of original images to be processed from the original images in the plurality of initial pairing data on the basis of the preset screening condition comprises:

determining original images to be processed of which the change angles of the parts to be adjusted are greater than the preset change angle threshold values in the original images;

wherein the parts to be adjusted comprise five sense organs.

5 . The method according to claim 1 , wherein processing the plurality of original images to be processed on the basis of the style model to be used to obtain style image to be used, comprises:

inputting each original image to be processed into the style model to be used to obtain the style image to be used corresponding to each original image,

wherein the style image to be used has different features from an initial style image corresponding to each original image to be processed.

6 . The method according to claim 1 , wherein performing deformation processing on the style image to be used to obtain the target style image corresponding to each original image to be processed, comprises:

determining pixel point information of key points in each original image to be processed and the style image to be used; and

determining deformation parameters on the basis of the pixel point information, and attaching the parts to be adjusted in each original image to be processed into the style image to be used on the basis of the deformation parameters, so as to obtain the target style image.

7 . The method according to claim 1 , wherein the method further comprises, before training the stylization conversion model to be trained on the basis of the stylized pairing data to obtain the target stylization conversion model:

determining a stylization conversion model to be trained of a target grid structure; and

splicing a discriminator to be trained for the stylization conversion model to be trained, and setting a parameter adjustment constraint condition for the discriminator to be trained to perform constraint adjustment on model parameters in the stylization conversion model to be trained and model parameters in the discriminator to be trained on the basis of the constraint condition, so as to obtain the target stylization conversion model.

8 . The method according to claim 7 , wherein training the stylization conversion model to be trained on the basis of the stylized pairing data to obtain the target stylization conversion model, comprises:

inputting each original image to be processed in the stylized pairing data into the stylization conversion model to be trained to obtain a second actual output image;

inputting the second actual output image and the target style image in the stylized pairing data into the discriminator to be trained to obtain a discrimination result;

adjusting model parameters in the stylization conversion model to be trained and model parameters in the discriminator to be trained on the basis of the discrimination result and the constraint condition; and

converging a loss function in the stylization conversion model to be trained and a loss function in the discriminator to be trained as training targets, so as to obtain the target stylization conversion model.

9 . The method according to claim 1 , further comprising:

deploying the target stylization conversion model in a client, and upon acquiring video frames to be processed, performing stylization processing on the video frames to be processed on the basis of the target stylization conversion model to obtain target video frames, and obtaining a target video on the basis of all target video frames.

10 . An electronic device, comprising:

a processor; and

a storage apparatus, configured to store a program which when executed by the processor, causes the processor to;

determine a plurality of initial pairing data, and performing training on the basis of the plurality of initial pairing data to obtain a style model to be used, wherein each initial pairing data comprises an original image and an initial style image obtained after the original image is processed by a three dimensional (3D) style generation model;

determine a plurality of original images to be processed from the original images in the plurality of initial pairing data on the basis of a preset screening condition, and processing each original image to be processed on the basis of the style model to be used, so as to obtain a style image to be used corresponding to each original image to be processed;

obtain a target style image corresponding to each original image to be processed by performing deformation processing on the style image to be used, and using each original image to be processed and the corresponding target style image as stylized pairing data; and

train a stylization conversion model to be trained on the basis of the stylized pairing data to obtain a target stylization conversion model, and upon acquiring video frames to be processed, performing stylization processing on the video frames to be processed on the basis of the target stylization conversion model to obtain a processed target video.

11 . The electronic device according to claim 10 , wherein the program causing the processor to determine the plurality of initial pairing data, causes the processor to:

acquire a plurality of original images comprising facial information; and

input each original image into a pre-trained 3D style generation model to obtain a corresponding initial style image after facial information processing.

12 . The electronic device according to claim 10 , wherein the program causing the processor to perform training on the basis of the plurality of initial pairing data to obtain the style model to be used, causes the processor to:

acquire a first style model to be trained;

for each initial pairing data, using an original image in each initial pairing data as an input of the first style model to be trained to obtain a first output image corresponding to the original image;

determine a loss value on the basis of the first output image and the initial style image corresponding to the original image, so as to adjust model parameters in the first style model to be trained on the basis of the loss value; and

converge a first loss function in the first style model to be trained as a training target to obtain the style model to be used.

13 . The electronic device according to claim 10 , wherein the preset screening condition comprises that change angles of parts to be adjusted are greater than preset change angle threshold values, and wherein the program causing the processor to determine the plurality of original images to be processed from the original images in the plurality of initial pairing data on the basis of the preset screening condition, causes the processor to:

determine original images to be processed of which the change angles of the parts to be adjusted are greater than the preset change angle threshold values in the original images;

wherein the parts to be adjusted comprise five sense organs.

14 . The electronic device according to claim 10 , wherein the program causing the processor to process the plurality of original images to be processed on the basis of the style model to be used to obtain style image to be used, causes the processor to:

input each original image to be processed into the style model to be used to obtain the style image to be used corresponding to each original image,

wherein the style image to be used has different features from an initial style image corresponding to each original image to be processed.

15 . The electronic device according to claim 10 , wherein the program causing the processor to perform deformation processing on the style image to be used to obtain the target style image corresponding to each original image to be processed, causes the processor to:

determine pixel point information of key points in each original image to be processed and the style image to be used; and

determine deformation parameters on the basis of the pixel point information, and attaching the parts to be adjusted in each original image to be processed into the style image to be used on the basis of the deformation parameters, so as to obtain the target style image.

16 . The electronic device according to claim 10 , wherein, before training the stylization conversion model to be trained on the basis of the stylized pairing data to obtain the target stylization conversion model, the program further causes the processor to:

determine a stylization conversion model to be trained of a target grid structure; and

splice a discriminator to be trained for the stylization conversion model to be trained, and setting a parameter adjustment constraint condition for the discriminator to be trained to perform constraint adjustment on model parameters in the stylization conversion model to be trained and model parameters in the discriminator to be trained on the basis of the constraint condition, so as to obtain the target stylization conversion model.

17 . The electronic device according to claim 16 , wherein the program causing the processor to train the stylization conversion model to be trained on the basis of the stylized pairing data to obtain the target stylization conversion model, causes the processor to:

input each original image to be processed in the stylized pairing data into the stylization conversion model to be trained to obtain a second actual output image;

input the second actual output image and the target style image in the stylized pairing data into the discriminator to be trained to obtain a discrimination result;

adjust model parameters in the stylization conversion model to be trained and model parameters in the discriminator to be trained on the basis of the discrimination result and the constraint condition; and

converge a loss function in the stylization conversion model to be trained and a loss function in the discriminator to be trained as training targets, so as to obtain the target stylization conversion model.

18 . The electronic device according to claim 10 , the program further causes the processor to:

deploy the target stylization conversion model in a client, and upon acquiring video frames to be processed, performing stylization processing on the video frames to be processed on the basis of the target stylization conversion model to obtain target video frames, and obtaining a target video on the basis of all target video frames.

19 . A non-transitory storage medium, comprising computer-executable instructions, wherein the computer-executable instructions are used for, when being executed by a computer processor, causes the computer processor to:

determine a plurality of initial pairing data, and performing training on the basis of the plurality of initial pairing data to obtain a style model to be used, wherein each initial pairing data comprises an original image and an initial style image obtained after the original image is processed by a three dimensional (3D) style generation model;

determine a plurality of original images to be processed from the original images in the plurality of initial pairing data on the basis of a preset screening condition, and processing each original image to be processed on the basis of the style model to be used, so as to obtain a style image to be used corresponding to each original image to be processed;

obtain a target style image corresponding to each original image to be processed by performing deformation processing on the style image to be used, and using each original image to be processed and the corresponding target style image as stylized pairing data; and

train a stylization conversion model to be trained on the basis of the stylized pairing data to obtain a target stylization conversion model, and upon acquiring video frames to be processed, performing stylization processing on the video frames to be processed on the basis of the target stylization conversion model to obtain a processed target video.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2026
From: LI, WENYUE
To: SHENZHEN JINRITOUTIAO TECHNOLOGY CO., LTD.
Reel/Frame 074684/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2026
From: ZHOU, CAIJIN
To: SHANGHAI GEWU ZHIYUAN NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 074684/0642 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2026
From: SHENZHEN JINRITOUTIAO TECHNOLOGY CO., LTD.; SHANGHAI GEWU ZHIYUAN NETWORK TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 074684/0791 →
Priority Claims (1)
CN 202210080456.8 · Jan 24, 2022 · national
Continuity (1)
Related Publication 20250173971A1 · May 29, 2025
References Cited (12)
US 8437514B2 · Wen · 2013 [cited by examiner]
US 10489683B1 · Koh · 2019 [cited by examiner]
US 11720994B2 · Luo · 2023 [cited by examiner]
US 20110098056A1 · Rhoads · 2011 [cited by examiner]
US 20180373999A1 · Xu · 2018 [cited by applicant]
US 20190340419A1 · Milman · 2019 [cited by examiner]
US 20230130535A1 · Ma · 2023 [cited by examiner]
US 20230164298A1 · Khot · 2023 [cited by examiner]
CN 113409342A · 2021 [cited by applicant]
CN 113658324A · 2021 [cited by applicant]
CN 114419300A · 2022 [cited by applicant]
ISA China National Intellectual Property Administration, International Search Report Issued in Application No. PCT/CN2023/072539, Mar. 21, 2023, WIPO, 3 pages. [cited by applicant]