IP Library › Granted Patent US 11,367,307
Granted Patent B2
US 11,367,307 · App. 17/105,186 · Granted Jun 21, 2022

Method for processing images and electronic device

Inventors: Shanshan Wu (Beijing, CN); Paliwan Pahaerding (Beijing, CN); Ni Ai (Beijing, CN)
Assignee: Beijing Dajia Internet Information Technology Co., Ltd.
G06V40/168A45D44/005G06T7/11G06T2207/20081G06T2207/20221G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,367,307
App. No.
17/105,186
Granted
Jun 21, 2022
Kind
B2
Abstract

Provided is a method for processing images. The method can include: acquiring a target face image, and performing face key point detection on the target face image; acquiring a first fusion image by fusing a virtual special effect and a face part matched in the target face image based on a face key point detection result; acquiring an occlusion mask of the target face image; and generating a second fusion image based on the occlusion mask and the first fusion image.

Claims (75)

1. A method for processing images, comprising:

acquiring a target face image, and performing face key point detection on the target face image;

acquiring a first fusion image by fusing a virtual special effect and a face part matched in the target face image based on a face key point detection result;

acquiring an occlusion mask of the target face image, wherein the occlusion mask is configured to indicate a face visible area which is not subject to an occluder and a face invisible area which is subject to the occluder in the target face image; and

generating a second fusion image based on the occlusion mask and the first fusion image.

2. The method according to claim 1 , wherein said acquiring the occlusion mask of the target face image comprises:

acquiring the face visible area and the face invisible area by semantically segmenting the target face image based on an image semantic segmentation model; and

generating the occlusion mask of the target face image, wherein in the occlusion mask, pixels taking a first value are configured to indicate the face visible area, and pixels taking a second value are configured to indicate the face invisible area.

3. The method according to claim 2 , wherein the image semantic segmentation model is trained by:

acquiring a training sample image and a label segmentation result of the training sample image, wherein the training sample image comprises an image in which a face area is subject to the occluder;

inputting the training sample image into a deep learning model;

determining, based on a target loss function, whether a predicted segmentation result of the training sample image output by the deep learning model matches the label segmentation result or not; and

acquiring the image semantic segmentation model by iteratively updating network parameters of the deep learning model until the deep learning model converges in response to the predicted segmentation result not matching the label segmentation result.

4. The method according to claim 1 , wherein said generating the second fusion image based on the occlusion mask and the first fusion image comprises:

determining the face visible area and the face invisible area in the first fusion image based on the occlusion mask;

retaining the virtual special effect in the face visible area of the first fusion image; and

acquiring the second fusion image by drawing the occluder on the face invisible area of the first fusion image in a fashion of being placed on a top layer.

5. The method according to claim 1 , wherein said acquiring the first fusion image by fusing the virtual special effect and the face part matched in the target face image based on the face key point detection result comprises:

determining the face part matching the virtual special effect;

fitting a target location area in the target face image based on the face key point detection result and the determined face part; and

acquiring the first fusion image by fusing the virtual special effect and the target location area of the target face image.

6. The method according to claim 1 , wherein said acquiring the target face image comprises:

performing face detection on an acquired image; and

acquiring the target face image by cropping a face area based on a face detection result in response to the acquired image comprising a face.

7. An electronic device, comprising:

a processor; and

a memory configured to store at least one computer program including at least one instruction executable by the processor;

wherein the at least one instruction, when executed by the processor, causes the processor to perform a method comprising:

acquiring a target face image, and performing face key point detection on the target face image;

acquiring a first fusion image by fusing a virtual special effect and a face part matched in the target face image based on a face key point detection result;

acquiring an occlusion mask of the target face image, wherein the occlusion mask is configured to indicate a face visible area which is not subject to an occluder and a face invisible area which is subject to the occluder in the target face image; and

generating a second fusion image based on the occlusion mask and the first fusion image.

8. The electronic device according to claim 7 , wherein said acquiring the occlusion mask of the target face image comprises:

acquiring the face visible area and the face invisible area by semantically segmenting the target face image based on an image semantic segmentation model; and

generating the occlusion mask of the target face image, wherein in the occlusion mask, pixels taking a first value are configured to indicate the face visible area, and pixels taking a second value are configured to indicate the face invisible area.

9. The electronic device according to claim 8 , wherein the image semantic segmentation model is trained by:

acquiring a training sample image and a label segmentation result of the training sample image, wherein the training sample image comprises an image in which a face area is subject to the occluder;

inputting the training sample image into a deep learning model;

determining, based on a target loss function, whether a predicted segmentation result of the training sample image output by the deep learning model matches the label segmentation result or not; and

acquiring the image semantic segmentation model by iteratively updating network parameters of the deep learning model until the deep learning model converges in response to the predicted segmentation result not matching the label segmentation result.

10. The electronic device according to claim 7 , wherein said generating the second fusion image based on the occlusion mask and the first fusion image comprises:

determining the face visible area and the face invisible area in the first fusion image based on the occlusion mask;

retaining the virtual special effect in the face visible area of the first fusion image; and

acquiring the second fusion image by drawing the occluder on the face invisible area of the first fusion image in a fashion of being placed on a top layer.

11. The electronic device according to claim 7 , wherein said acquiring the first fusion image by fusing the virtual special effect and the face part matched in the target face image based on the face key point detection result comprises:

determining the face part matching the virtual special effect;

fitting a target location area in the target face image based on the face key point detection result and the determined face part; and

acquiring the first fusion image by fusing the virtual special effect and the target location area of the target face image.

12. The electronic device according to claim 7 , wherein said acquiring the target face image comprises:

performing face detection on an acquired image; and

acquiring the target face image by cropping a face area based on a face detection result in response to the acquired image comprising a face.

13. A non-transitory computer-readable storage medium storing at least one computer program including at least one instruction, wherein the at least one instruction, when executed by a processor of an electronic device, causes the electronic device to perform a method comprising:

acquiring a target face image, and performing face key point detection on the target face image;

acquiring a first fusion image by fusing a virtual special effect and a face part matched in the target face image based on a face key point detection result;

acquiring an occlusion mask of the target face image, wherein the occlusion mask is configured to indicate a face visible area which is not subject to an occluder and a face invisible area which is subject to the occluder in the target face image; and

generating a second fusion image based on the occlusion mask and the first fusion image.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein said acquiring the occlusion mask of the target face image comprises:

acquiring the face visible area and the face invisible area by semantically segmenting the target face image based on an image semantic segmentation model; and

generating the occlusion mask of the target face image, wherein in the occlusion mask, pixels taking a first value are configured to indicate the face visible area, and pixels taking a second value are configured to indicate the face invisible area.

15. The non-transitory computer-readable storage medium according to claim 14 , wherein the image semantic segmentation model is trained by:

acquiring a training sample image and a label segmentation result of the training sample image, wherein the training sample image comprises an image in which a face area is subject to the occluder;

inputting the training sample image into a deep learning model;

determining, based on a target loss function, whether a predicted segmentation result of the training sample image output by the deep learning model matches the label segmentation result or not; and

acquiring the image semantic segmentation model by iteratively updating network parameters of the deep learning model until the deep learning model converges in response to the predicted segmentation result not matching the label segmentation result.

16. The non-transitory computer-readable storage medium according to claim 13 , wherein said generating the second fusion image based on the occlusion mask and the first fusion image comprises:

determining the face visible area and the face invisible area in the first fusion image based on the occlusion mask;

retaining the virtual special effect in the face visible area of the first fusion image; and

acquiring the second fusion image by drawing the occluder on the face invisible area of the first fusion image in a fashion of being placed on a top layer.

17. The non-transitory computer-readable storage medium according to claim 13 , wherein said acquiring the first fusion image by fusing the virtual special effect and the face part matched in the target face image based on the face key point detection result comprises:

determining the face part matching the virtual special effect;

fitting a target location area in the target face image based on the face key point detection result and the determined face part; and

acquiring the first fusion image by fusing the virtual special effect and the target location area of the target face image.

18. The non-transitory computer-readable storage medium according to claim 13 , wherein said acquiring the target face image comprises:

performing face detection on an acquired image; and

acquiring the target face image by cropping a face area based on a face detection result in response to the acquired image comprising a face.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2021
From: WU, SHANSHAN; AI, NI; PAHAERDING, PALIWAN
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 055236/0678 →
Priority Claims (1)
CN 201911168483.5 · Nov 25, 2019 · national
Continuity (1)
Related Publication 20210158021A1 · May 27, 2021