IP Library Granted Patent US 12705925
Granted Patent B2
US 12705925 · App. 18/571,579 · Granted Aug 11, 2026

Method, electronic device, and storage medium for image processing

Inventor: Ziyang Cheng (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
G06V40/16G06V10/77
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705925
App. No.
18/571,579
Granted
Aug 11, 2026
Kind
B2
Abstract

Embodiments of the disclosure provide a method, apparatus, electronic device ( 700 ), and storage medium for image processing. The method inputting a to-be-processed facial image to a predetermined model (S 110 ); and outputting, by the predetermined model, a target facial image (S 120 ) with a predetermined object removed from the to-be-processed facial image; wherein the predetermined model is trained and generated based on an attention map (a) of the predetermined object. Since the predetermined model is trained based on the attention map (a) of the predetermined object, it is able to first generate the attention map (a) of the predetermined object based on unpaired data training, and then train to remove the predetermined object from the facial image with the attention map (a) of the predetermined object.

Claims (67)

1 . A method of image processing, comprising:

inputting a to-be-processed facial image to a predetermined model; and

outputting, by the predetermined model, a target facial image with a predetermined object removed from the to-be-processed facial image;

wherein the predetermined model is trained and generated based on an attention map of the predetermined object; and

wherein the predetermined model is generated based on the following:

training a first model based on a first facial image containing the predetermined object and a second facial image without the predetermined object;

outputting, by the trained first model, an attention map of the predetermined object in the first facial image; and

training a second model based on the first facial image and the attention map; and

generating the predetermined model based on the trained first model and the trained second model.

2 . The method of image processing according to claim 1 , wherein the predetermined object comprises a beard, a bang, or an eye bag.

3 . The method of image processing according to claim 1 , wherein the first model is trained based on the following:

setting different image labels for the first facial image and the second facial image;

inputting the first facial image, the second facial image, and the image labels corresponding to the respective facial images into the first model;

determining a candidate object by the first model, and outputting prediction labels for the first facial image and the second facial image based on the candidate object; and

training the first model based on the prediction labels and the set image labels, and determining the candidate object determined by the trained first model as the predetermined object.

4 . The method of image processing according to claim 1 , wherein the second model is trained based on the following:

inputting the first facial image and the attention map to the second model, and outputting, by the second model, a third facial image with the predetermined object removed from the first facial image; and

inputting the second facial image and the third facial image to a first discriminator and training the second model based on a result of the first discriminator.

5 . The method of image processing according to claim 3 , wherein outputting, by the second model, a third facial image with the predetermined object removed from the first facial image comprises:

processing, by the second model and based on the attention map, pixel points corresponding to the predetermined object in the first facial image, and outputting the third facial image with the predetermined object removed.

6 . The method of image processing according to claim 5 , wherein processing pixel points corresponding to the predetermined object in the first facial image comprises:

copying and transferring a pixel point not labeled by the attention map in the first facial image to a location of a pixel point labeled by the attention map; and

wherein the pixel point labeled by the attention map belongs to the predetermined object.

7 . The method of image processing according to claim 5 , further comprising: before outputting the third facial image with the predetermined object removed, performing predetermined adjusting processing the third facial image.

8 . The method of image processing according to claim 1 , wherein generating the predetermined model based on the trained first model and the trained second model comprises:

establishing a connection between an output layer of the trained first model and an input layer of the trained second model, to integrate into the predetermined model.

9 . The method of claim 1 , wherein the first facial image is obtained based on the following:

obtaining a first number of fourth facial images containing the predetermined object and fifth facial images that are corresponding to the fourth facial images and contain no predetermined object, and a second number of sixth facial images containing the predetermined object; wherein the second number is greater than the first number;

pre-training the third model based on the fourth facial images and the fifth facial images; inputting the sixth facial images into the pre-trained third model, and determining an image output by the pre-trained third model as the first facial image; and

wherein generating the predetermined model based on the trained first model and the trained second model comprises:

processing the first facial image, the trained first model, and the trained second model to obtain a third facial image with the predetermined object removed from the first facial image; and optimizing and training the third model based on the third facial image and the sixth facial image, and determining the optimized and trained third model as the predetermined model.

10 . The method of claim 9 , wherein the third model is pre-trained based on the following:

inputting the fourth facial image to the third model, to cause the third model to output a seventh facial image; and

inputting the fifth facial image and the seventh facial image to a second discriminator, and pre-training the third model based on a result of the second discriminator.

11 . An electronic device, comprising: one or more processors; and

a storage apparatus configured to store one or more programs;

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement the method of image processing comprising:

inputting a to-be-processed facial image to a predetermined model; and

outputting, by the predetermined model, a target facial image with a predetermined object removed from the to-be-processed facial image;

wherein the predetermined model is trained and generated based on an attention map of the predetermined object; and

wherein the predetermined model is generated based on the following:

training a first model based on a first facial image containing the predetermined object and a second facial image without the predetermined object;

outputting, by the trained first model, an attention map of the predetermined object in the first facial image; and

training a second model based on the first facial image and the attention map; and generating the predetermined model based on the trained first model and the trained second model.

12 . The electronic device of claim 11 , wherein the predetermined object comprises a beard, a bang, or an eye bag.

13 . The electronic device of claim 11 , wherein the first model is trained based on the following:

setting different image labels for the first facial image and the second facial image;

inputting the first facial image, the second facial image, and the image labels corresponding to the respective facial images into the first model;

determining a candidate object by the first model, and outputting prediction labels for the first facial image and the second facial image based on the candidate object; and

training the first model based on the prediction labels and the set image labels, and determining the candidate object determined by the trained first model as the predetermined object.

14 . The electronic device of claim 11 , wherein the second model is trained based on the following:

inputting the first facial image and the attention map to the second model, and outputting, by the second model, a third facial image with the predetermined object removed from the first facial image; and

inputting the second facial image and the third facial image to a first discriminator and training the second model based on a result of the first discriminator.

15 . The electronic device of claim 14 , wherein outputting, by the second model, a third facial image with the predetermined object removed from the first facial image comprises:

processing, by the second model and based on the attention map, pixel points corresponding to the predetermined object in the first facial image, and outputting the third facial image with the predetermined object removed.

16 . The electronic device of claim 15 , wherein processing pixel points corresponding to the predetermined object in the first facial image comprises:

copying and transferring a pixel point not labeled by the attention map in the first facial image to a location of a pixel point labeled by the attention map; and

wherein the pixel point labeled by the attention map belongs to the predetermined object.

17 . The electronic device of claim 15 , wherein the method further comprises, before outputting the third facial image with the predetermined object removed, performing predetermined adjusting processing the third facial image.

18 . A non-transitory storage medium comprising computer-executable instructions, the computer-executable instructions, when executed by a computer processor, causing the method of image processing comprising:

inputting a to-be-processed facial image to a predetermined model; and

outputting, by the predetermined model, a target facial image with a predetermined object removed from the to-be-processed facial image;

wherein the predetermined model is trained and generated based on an attention map of the predetermined object; and

wherein the predetermined model is generated based on the following:

training a first model based on a first facial image containing the predetermined object and a second facial image without the predetermined object;

outputting, by the trained first model, an attention map of the predetermined object in the first facial image; and

training a second model based on the first facial image and the attention map; and generating the predetermined model based on the trained first model and the trained second model.