IP Library › Granted Patent US 11,403,874
Granted Patent B2
US 11,403,874 · App. 16/994,148 · Granted Aug 2, 2022

Virtual avatar generation method and apparatus for generating virtual avatar including user selected face property, and storage medium

Inventors: Tinghao Liu (Beijing, CN); Lichen Zhao (Beijing, CN); Quan Wang (Beijing, CN); Chen Qian (Beijing, CN)
Assignee: BEIJING SENSETIME TECHNOLOGY DEVELOPMENT CO., LTD.
G06V40/168G06T13/40G06V30/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,403,874
App. No.
16/994,148
Filed
Aug 14, 2020
Granted
Aug 2, 2022
Kind
B2
Examiner
NGUYEN, VU
Art Unit
2619
USPC
148/421
Abstract

A virtual avatar generation method includes: determining a target task associated with at least one target face property, where the at least one target face property is one of a plurality of predefined face properties respectively; performing, according to the target task, target face property analysis on a target image including at least a face to obtain a target face property feature associated with the target face property of the target image; determining a target virtual avatar template corresponding to the target face property feature according to predefined correspondence between face property features and virtual avatar templates; and generating a virtual avatar of the target image based on the target virtual avatar template.

Claims (54)

1. A virtual avatar generation method, comprising:

determining a target task associated with at least one target face property, wherein the target face property is one of a plurality of predefined face properties, the target face property comprises at least one predefined subclass, and each of the at least one predefined subclass comprises at least one face property feature;

performing, according to the target task, target face property analysis on a target image comprising at least a face, to obtain a target face property feature associated with the target face property of the target image;

determining a target virtual avatar template corresponding to the target face property feature according to predefined correspondence between face property features and virtual avatar templates; and

generating a virtual avatar of the target image based on the target virtual avatar template;

wherein performing, according to the target task, target face property analysis on a target image comprising at least a face, to obtain a target face property feature associated with the target face property of the target image comprises:

determining a target neural network corresponding to the target face property;

inputting the target image into the target neural network to obtain estimated values output from the target neural network, wherein the estimated values represent respective probabilities that the target image has one or more face property features associated with the target face property; and

for a first subclass of the at least one subclass included in the target face property, taking, as the target face property feature corresponding to the first subclass, a face property feature corresponding to a maximum value among the estimated values output from the target neural network for the first subclass.

2. The method according to claim 1 , wherein the target neural network is trained by:

inputting at least one sample image comprising at least a face into a first neural network, wherein each of the at least one sample image is labeled with at least one face property feature that is associated with a first face property of the plurality of predefined face properties, and the first neural network comprises a first sub-network corresponding to the first face property; and

training the first sub-network by taking at least one face property feature which is output from the first neural network and is associated with the first face property of the at least one sample image as a predicted value, and taking the at least one face property feature which is labeled on the at least one sample image and corresponds to the first face property as a real value, so as to obtain the target neural network after the training.

3. The method according to claim 2 , wherein the first sub-network has a network structure of a residual neural network and comprises at least one residual unit.

4. The method according to claim 3 , wherein

each of the at least one residual unit comprises at least one convolutional layer and at least one batch normalization layer; and

in a case where the at least one residual unit comprises a plurality of residual units, a number of convolutional layers and a number of batch normalization layers included in a second residual unit of the plurality of residual units both are greater than that included in a first residual unit of the plurality of residual units.

5. The method according to claim 3 , wherein

the first sub-network further comprises an output segmentation layer, and

the output segmentation layer is configured to segment, according to one or more predefined subclasses included in the first face property, feature information extracted from the sample image, to obtain respective estimated values for one or more face property features respectively associated with the one or more subclasses.

6. The method according to claim 2 , further comprising:

performing affine transformation on an image of interest to obtain a frontalized face image; and

clipping an image of a target region from the frontalized face image to obtain the target image or the sample image, wherein the target region comprises at least a region where a face key point is located.

7. The method according to claim 6 , wherein the target region further comprises a region with a preset area outside a face part corresponding to the target face property.

8. A virtual avatar generation apparatus, comprising:

a processor; and

a memory storing instructions executable by the processor,

wherein the processor is configured to execute the instructions stored in the memory to perform operations comprising:

determining a target task associated with at least one target face property, wherein the target face property is one of a plurality of predefined face properties, the target face property comprises at least one predefined subclass, and each of the at least one predefined subclass comprises at least one face property feature;

determining a target neural network corresponding to the target face property;

inputting a target image comprising at least a face into the target neural network to obtain estimated values output from the target neural network, wherein the estimated values represent respective probabilities that the target image has one or more face property features associated with the target face property; and

for a first subclass of the at least one subclass included in the target face property, taking, as the target face property feature corresponding to the first subclass, a face property feature corresponding to a maximum value among the estimated values output from the target neural network for the first subclass;

determining a target virtual avatar template corresponding to the target face property feature according to predefined correspondence between face property features and virtual avatar templates; and

generating a virtual avatar of the target image based on the target virtual avatar template.

9. The apparatus according to claim 8 , wherein the operations further comprise:

inputting at least one sample image comprising at least a face into a first neural network, wherein each of the at least one sample image is labeled with at least one face property feature that is associated with a first face property of the plurality of predefined face properties, and the first neural network comprises a first sub-network corresponding to the first face property; and

training the first sub-network by taking at least one face property feature which is output from the first neural network and is associated with the first face property of the at least one sample image as a predicted value, and taking the at least one face property feature which is labeled on the at least one sample image and corresponds to the first face property as a real value, so as to obtain the target neural network after the training.

10. The apparatus according to claim 9 , wherein the first sub-network has a network structure of a residual neural network and comprises at least one residual unit.

11. The apparatus according to claim 10 , wherein

each of the at least one residual unit comprises at least one convolutional layer and at least one batch normalization layer; and

in a case where the at least one residual unit comprises a plurality of residual units, a number of convolutional layers and a number of batch normalization layers included in a second residual unit of the plurality of residual units both are greater than that included in a first residual unit of the plurality of residual units.

12. The apparatus according to claim 10 , wherein

the first sub-network further comprises an output segmentation layer, and

the output segmentation layer is configured to segment, according to one or more predefined subclasses included in the first face property, feature information extracted from the sample image, to obtain respective estimated values for one or more face property features respectively associated with the one or more subclasses.

13. The apparatus according to claim 9 , wherein the operations further comprise:

performing affine transformation on an image of interest to obtain a frontalized face image; and

clipping an image of a target region from the frontalized face image to obtain the target image or the sample image, wherein the target region comprises at least a region where a face key point is located.

14. The apparatus according to claim 13 , wherein the target region further comprises a region with a preset area outside a face part corresponding to the target face property.

15. A non-transitory computer-readable storage medium storing a computer program which, when executed by one or more processors, causes the one or more processors to perform operations comprising:

determining a target task associated with at least one target face property, wherein the target face property is one of a plurality of predefined face properties, the target face property comprises at least one predefined subclass, and each of the at least one predefined subclass comprises at least one face property feature;

determining a target neural network corresponding to the target face property;

inputting a target image comprising at least a face into the target neural network to obtain estimated values output from the target neural network, wherein the estimated values represent respective probabilities that the target image has one or more face property features associated with the target face property; and

for a first subclass of the at least one subclass included in the target face property, taking, as the target face property feature corresponding to the first subclass, a face property feature corresponding to a maximum value among the estimated values output from the target neural network for the first subclass;

determining a target virtual avatar template corresponding to the target face property feature according to predefined correspondence between face property features and virtual avatar templates; and

generating a virtual avatar of the target image based on the target virtual avatar template.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2020
From: LIU, TINGHAO; ZHAO, LICHEN; WANG, QUAN; QIAN, CHEN
To: BEIJING SENSETIME TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 053503/0373 →
Priority Claims (1)
CN 201910403642.9 · May 15, 2019 · national
Continuity (2)
Continuation PCTCN2020074597 · Feb 10, 2020
Related Publication 20200380246A1 · Dec 3, 2020
Cited By (1)
US 12,586,285