IP Library › Granted Patent US 12,499,514
Granted Patent B2
US 12,499,514 · App. 18/252,855 · Granted Dec 16, 2025

Animal face style image generation method and apparatus, model training method and apparatus, and device

Inventor: Qian He (Beijing, CN)
Assignee: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
G06T5/50G06T7/12G06T11/60G06V40/171G06T2207/20081G06T2207/20221G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,514
App. No.
18/252,855
Filed
May 12, 2023
Granted
Dec 16, 2025
Kind
B2
Examiner
WU, YANNA
Art Unit
2615
USPC
345/629
Abstract

An animal face style image generation method, a model training method, and a device are provided. The method includes: acquiring an original human face image; and obtaining an animal face style image corresponding to the original human face image. The animal face style image refers to an image obtained by transforming a human face on the original human face image into an animal face, the animal face style image generation model is obtained by training with a first human face sample image and a first animal face style sample image, the first animal face style sample image is generated from the first human face sample image by a pre-trained animal face generation model, and the animal face generation model is obtained by training with a second human face sample image and a first animal face sample image.

Claims (73)

1 . A method for generating an animal face style image, implemented by an electronic device with a display interface, a processor and a storage, comprising:

acquiring, by the processor, an original human face image, from the storage of the electronic device or from a camera capturing the original human face image;

obtaining, by the processor, an animal face style image corresponding to the original human face image by using a pre-trained animal face style image generation model;

wherein the animal face style image is obtained by transforming a human face on the original human face image into an animal face, the animal face style image generation model is obtained by training with a first human face sample image and a first animal face style sample image, the first animal face style sample image is generated from the first human face sample image by a pre-trained animal face generation model, and the animal face generation model is obtained by training with a second human face sample image and a first animal face sample image; and

rendering and displaying the animal face style image on the display interface of the electronic device.

2 . The method according to claim 1 , further comprising:

determining, based on an animal special effect type selected by a user, a correspondence between animal face key points and human face key points, wherein the correspondence corresponds to the animal special effect type; and

performing, based on the correspondence between the animal face key points and the human face key points, human face position adjustment on a user image to obtain the original human face image, wherein the correspondence corresponds to the animal special effect type, and the original human face image meets input requirements of the animal face style image generation model.

3 . The method according to claim 2 , further comprising:

fusing an animal face region in the animal face style image with a background region in the user image to obtain a target animal face style image corresponding to the user image.

4 . The method according to claim 3 , wherein the fusing the animal face region in the animal face style image with the background region in the user image to obtain the target animal face style image corresponding to the user image comprises:

obtaining, based on the animal face style image, an intermediate image with a same image size as the user image, wherein a position of an animal face region in the intermediate image is the same as a position of a human face region in the user image;

determining a first animal face mask image corresponding to the animal special effect type; and

fusing, based on the first animal face mask image, the user image with the intermediate image to obtain the target animal face style image corresponding to the user image, wherein the first animal face mask image is used for determining the animal face region in the intermediate image as an animal face region in the target animal face style image.

5 . The method according to claim 1 , wherein:

the first human face sample image is obtained by performing, based on a first correspondence between face key points in a first original human face sample image and animal face key points in a first original animal face sample image, human face position adjustment on the first original human face sample image;

the second human face sample image is obtained by performing, based on a second correspondence between face key points in a second original human face sample image and animal face key points in the first original animal face sample image, human face position adjustment on the second original human face sample image; and

the first animal face sample image is obtained by performing, based on the first correspondence or the second correspondence, animal face position adjustment on the first original animal face sample image.

6 . The method according to claim 1 , wherein:

the animal face style image generation model is obtained by training with the first human face sample image and a second animal face style sample image, and the second animal face style sample image is obtained by replacing a background region in the first animal face style sample image with a background region in the first human face sample image.

7 . The method according to claim 6 , wherein:

the second animal face style sample image is obtained by fusing, based on a second animal face mask image, the first animal face style sample image with the first human face sample image; and

the second animal face mask image is obtained from the first animal face style sample image by a pre-trained animal face segmentation model, wherein the second animal face mask image is used for determining an animal face region in the first animal face style sample image as an animal face region in the second animal face style sample image.

8 . A method for training an animal face style image generation model, implemented by an electronic device with a display interface, a processor and a storage, comprising:

training, by the processor, an image generation model with a second human face sample image and a first animal face sample image to obtain an animal face generation model, wherein the second human face sample image and the first animal face sample image is obtained from the storage of the electronic device;

obtaining, by the processor, a first animal face style sample image corresponding to a first human face sample image by using the animal face generation model, wherein the first animal face style sample image refers to an image obtained by transforming a human face on the first human face sample image into an animal face; and

training, by the processor, a style image generation model with the first human face sample image and the first animal face style sample image to obtain an animal face style image generation model;

wherein, the animal face style image generation model is used for obtaining an animal face style image corresponding to an original human face image by transforming a human face on the original human face image into an animal face, and the animal face style image is to be rendered and displayed on the display interface of the electronic device.

9 . The method according to claim 8 , further comprising:

determining a second correspondence between human face key points in a second original human face sample image and animal face key points in a first original animal face sample image;

performing, based on the second correspondence, human face position adjustment on the second original human face sample image to obtain the second human face sample image; and

performing, based on the second correspondence, animal face position adjustment on the first original animal face sample image based on the second correspondence to obtain the first animal face sample image.

10 . The method according to claim 9 , further comprising:

determining a first correspondence between face key points in a first original human face sample image and animal face key points in the first original animal face sample image; and

performing, based on the first correspondence, animal face position adjustment on the first original human face sample image to obtain the first human face sample image.

11 . The method according to claim 8 , further comprising:

replacing a background region in the first animal face style sample image with a background region in the first human face sample image to obtain a second animal face style sample image; and

the training the style image generation model with the first human face sample image and the first animal face style sample image to obtain the animal face style image generation model comprises:

training the style image generation model with the first human face sample image and a second animal face style sample image to obtain the animal face style image generation model.

12 . The method according to claim 11 , wherein the replacing the background region in the first animal face style sample image with the background region in the first human face sample image to obtain the second animal face style sample image comprises:

obtaining an animal face mask image corresponding to the first animal face style sample image by using a pre-trained animal face segmentation model; and

fusing, based on the animal face mask image, the first animal face style sample image with the first human face sample image to obtain the second animal face style sample image, wherein the animal face mask image is used for determining an animal face region in the first animal face style sample image as an animal face region in the second animal face style sample image.

13 . The method according to claim 12 , further comprising:

acquiring a second animal face sample image and a marking of position of the animal face region in the second animal face sample image; and

obtaining the animal face segmentation model by training with the second animal face sample image and the marking of position of the animal face region.

14 . An electronic device, comprising:

a display interface;

a processor; and

a memory storing a computer program that, when executed by the processor, causes the processor to implement:

acquiring an original human face image, from a storage of the electronic device or from a camera capturing the original human face image;

obtaining an animal face style image corresponding to the original human face image by using a pre-trained animal face style image generation model;

wherein the animal face style image is obtained by transforming a human face on the original human face image into an animal face, the animal face style image generation model is obtained by training with a first human face sample image and a first animal face style sample image, the first animal face style sample image is generated from the first human face sample image by a pre-trained animal face generation model, and the animal face generation model is obtained by training with a second human face sample image and a first animal face sample image; and

rendering and displaying the obtained animal face style image on the display interface of the electronic device.

15 . The electronic device according to claim 14 , wherein the processor further implements:

determining, based on an animal special effect type selected by a user, a correspondence between animal face key points and human face key points, wherein the correspondence corresponds to the animal special effect type; and

performing, based on the correspondence between the animal face key points and the human face key points, human face position adjustment on a user image to obtain the original human face image, wherein the correspondence corresponds to the animal special effect type, and the original human face image meets input requirements of the animal face style image generation model.

16 . The electronic device according to claim 15 , wherein the processor further implements:

fusing an animal face region in the animal face style image with a background region in the user image to obtain a target animal face style image corresponding to the user image.

17 . An electronic device, comprising:

a display interface;

a processor; and

a memory storing a computer program that, when executed by the processor, causes the processor to implement:

training an image generation model with a second human face sample image and a first animal face sample image to obtain an animal face generation model, wherein the second human face sample image and the first animal face sample image is obtained from the storage of the electronic device;

obtaining a first animal face style sample image corresponding to a first human face sample image by using the animal face generation model, wherein the first animal face style sample image refers to an image obtained by transforming a human face on the first human face sample image into an animal face; and

training a style image generation model with the first human face sample image and the first animal face style sample image to obtain an animal face style image generation model;

wherein, the animal face style image generation model is used for obtaining an animal face style image corresponding to an original human face image by transforming a human face on the original human face image into an animal face, and the animal face style image is to be rendered and displayed on the display interface of the electronic device.

18 . The electronic device according to claim 17 , wherein the processor further implements:

determining a second correspondence between human face key points in a second original human face sample image and animal face key points in a first original animal face sample image;

performing, based on the second correspondence, human face position adjustment on the second original human face sample image to obtain the second human face sample image; and

performing, based on the second correspondence, animal face position adjustment on the first original animal face sample image based on the second correspondence to obtain the first animal face sample image.

19 . The electronic device according to claim 18 , wherein the processor further implements:

determining a first correspondence between face key points in a first original human face sample image and animal face key points in the first original animal face sample image; and

performing, based on the first correspondence, animal face position adjustment on the first original human face sample image to obtain the first human face sample image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: HE, QIAN
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 063631/0074 →
Priority Claims (1)
CN 202011269334.0 · Nov 13, 2020 · national
Continuity (1)
Related Publication 20240005466A1 · Jan 4, 2024
References Cited (13)
US 20180189553A1 · Guo et al. · 2018 [cited by applicant]
US 20210334942A1 · Wang · 2021 [cited by examiner]
CN 110930297A · 2020 [cited by applicant]
CN 111783647A · 2020 [cited by applicant]
CN 111968029A · 2020 [cited by applicant]
CN 112330534A · 2021 [cited by applicant]
CN 112989904A · 2021 [cited by applicant]
WO 2019198850A1 · 2019 [cited by applicant]
International Search Report (with English translation) and Written Opinion issued in PCT/CN2021/130301, dated Feb. 10, 2022, 14 pages provided. [cited by applicant]
Choi et al., “StarGAN v2: Diverse Image Synthesis for Multiple Domains”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Aug. 5, 2020, pp. 8185-8194. [cited by applicant]
Liang Zi Wei, “Generating the “cat and dog version” of Trump, and breaking the face editing tool StarGANv2” , https://ishare.ifeng.com/c/s/7w2keWci1Dq, Apr. 28, 2020, with English translation. [cited by applicant]
Office Action issued in Japanese Application No. 2023-528414, dated Apr. 23, 2024, with machine translation. [cited by applicant]
Decision to Grant a Patent for Japanese Patent Application No. 2023-528414, mailed on Oct. 29, 2024, 5 pages. [cited by applicant]