IP Library › Granted Patent US 12,548,210
Granted Patent B2
US 12,548,210 · App. 18/584,777 · Granted Feb 10, 2026

Local attribute image editing using an image generation model and a feature image generation model

Inventors: Haokun Chen (Shenzhen, CN); Ruixue Shen (Shenzhen, CN); Rui Wang (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06T11/00G06T3/40G06T7/337G06T2207/20081G06T2207/20092G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,210
App. No.
18/584,777
Granted
Feb 10, 2026
Kind
B2
Abstract

An image editing method includes acquiring an initial image generation model and a feature image generation model, the initial image generation model having been trained based on a first training image set, the feature image generation model having been obtained by training the initial image generation model based on a second training image set. The method further includes acquiring a joint mask image based on image regions corresponding to the target attribute in the object images, and acquiring a second initial object image and a second feature object image output by corresponding target network layers. The method further includes fusing the second initial object image and the second feature object image based on the joint mask image to obtain a reference object image.

Claims (95)

1 . An image editing method, comprising:

acquiring an initial image generation model and a feature image generation model, the initial image generation model having been trained based on a first training image set, the feature image generation model having been obtained by training the initial image generation model based on a second training image set, object images in the first training image set and the second training image set including objects of a same category, each object image in the second training image set comprising a target attribute, and an image output by the feature image generation model having the target attribute;

based on an image editing request, inputting a value of a first variable as input data into the initial image generation model and the feature image generation model separately;

in response to the input data, acquiring an image output by the initial image generation model to obtain a first initial object image of the category, and acquiring an image output by the feature image generation model to obtain a first feature object image of the category;

based on image regions corresponding to the target attribute in the object images, acquiring attribute mask images respectively corresponding to the first initial object image and the first feature object image, and obtaining a joint mask image based on the attribute mask images;

acquiring images output by corresponding target network layers in the initial image generation model and the feature image generation model to obtain a second initial object image and a second feature object image; and

fusing the second initial object image and the second feature object image based on the joint mask image to obtain a reference object image, as a result of performing target attribute editing on the second initial object image.

2 . The method according to claim 1 , wherein the obtaining the initial image generation model and the feature image generation model comprises:

performing adversarial learning on an initial image generation network and an initial image discrimination network based on the first training image set to obtain an intermediate image generation network and an intermediate image discrimination network;

obtaining the initial image generation model based on the intermediate image generation network;

performing adversarial learning on the intermediate image generation network and the intermediate image discrimination network based on the second training image set to obtain a target image generation network; and

obtaining the feature image generation model based on the target image generation network.

3 . The method according to claim 1 , the method further comprising:

acquiring a first candidate image set, first candidate images in the first candidate image set being object images corresponding to objects of the category;

performing image alignment on the first candidate images based on positions of a reference object part of the objects of the category in the first candidate images;

obtaining the first training image set based on the first candidate images subjected to the image alignment;

acquiring a second candidate image set, second candidate images in the second candidate image set being object images corresponding to the objects of the category and comprising the target attribute;

performing image alignment on the second candidate images based on positions of the reference object part of the objects of the category in the second candidate images; and

obtaining the second training image set based on the second candidate images subjected to the image alignment.

4 . The method according to claim 1 , wherein the obtaining the target joint mask image based on the attribute mask images comprises:

using an attribute mask image corresponding to the first initial object image as a first mask image, and using an attribute mask image corresponding to the first feature object image as a second mask image, the first initial object image and the first feature object image having different sizes, and the first mask image and the second mask image having different sizes; and

performing size alignment on the first mask image and the second mask image, and obtaining the joint mask image based on the first mask image and the second mask image subjected to the size alignment.

5 . The method according to claim 1 , wherein the fusing the second initial object image and the second feature object image comprises:

acquiring, from the second initial object image, a first image region matching a shielded region in the joint mask image;

fusing the second initial object image and the second feature object image to obtain a fused object image;

acquiring, from the fused object image, a second image region matching a non-shielded region in the joint mask image; and

obtaining the reference object image based on the first image region and the second image region.

6 . The method according to claim 5 , wherein

the acquiring, from the second initial object image, the first image region matching a shielded region in the target joint mask image comprises:

performing size transformation on the joint mask image to obtain a transformed joint mask image of a same size as the second initial object image;

performing re-masking processing on the transformed joint mask image to obtain a re-masked joint mask image; and

fusing the second initial object image and the re-masked joint mask image to obtain the first image region; and

the acquiring, from the fused object image, the second image region matching a non-shielded region in the joint mask image comprises:

fusing the fused object image and the transformed joint mask image to obtain the second image region.

7 . The method according to claim 1 , the method further comprising:

replacing the second initial object image with the reference object image and inputting the reference object image into downstream network layers of the target network layer in the initial image generation model, and acquiring an image output by an end network layer of the initial image generation model as a target object image, wherein

the target object image is a result of performing target attribute editing on an original object image, which is output by the end network layer of the initial image generation model after the value of the first variable is input into the initial image generation model.

8 . The method according to claim 7 , wherein the replacing the second initial object image with the reference object image comprises:

replacing the second initial object image with the reference object image and inputting the reference object image into the downstream network layers of the target network layer in the initial image generation model, and acquiring an image output by a third network layer in the downstream network layers as a third initial object image;

acquiring an image output by a fourth network layer corresponding to the third network layer in the feature image generation model as a third feature object image;

fusing the third initial object image and the third feature object image based on a current joint mask image to obtain an updated object image, the current joint mask image being the joint mask image or an updated joint mask image obtained based on an image output by a fifth network layer in the initial image generation model and an image output by a sixth network layer in the feature image generation model; and

replacing the third initial object image with the updated object image and inputting the updated object image into downstream network layers of the third network layer in the initial image generation model, and acquiring an image output by the end network layer of the initial image generation model as the target object image.

9 . The method according to claim 7 , the method further comprising:

using the original object image and the target object image as a training image pair; and

performing model training on an initial image attribute editing model based on the training image pair to obtain a trained image attribute editing model, wherein the trained image attribute editing model is configured to perform target attribute editing on an input image.

10 . The method according to claim 1 , wherein network layers connected in sequence in each of the initial image generation model and the feature image generation model are configured to output images having gradually increased sizes, and in the initial image generation model and the feature image generation model, images output by corresponding network layers in have a same size.

11 . The method according to claim 1 , wherein the objects of the category are faces, the initial image generation model is an initial face image generation model, the feature image generation model is a feature face image generation model, and the target attribute is a local face attribute.

12 . An image editing apparatus, the apparatus comprising:

processing circuitry configured to

acquire an initial image generation model and a feature image generation model, the initial image generation model having been trained based on a first training image set, the feature image generation model having been obtained by training the initial image generation model based on a second training image set, object images in the first training image set and the second training image set including objects of a same category, each object image in the second training image set comprising a target attribute, and an image output by the feature image generation model having the target attribute;

based on an image editing request, input a value of a first variable as input data into the initial image generation model and the feature image generation model separately;

in response to the input data, acquire an image output by the initial image generation model to obtain a first initial object image of the category, and acquire an image output by the feature image generation model to obtain a first feature object image of the category;

based on image regions corresponding to the target attribute in the object images, acquire attribute mask images respectively corresponding to the first initial object image and the first feature object image, and obtain a joint mask image based on the attribute mask images;

acquire images output by corresponding target network layers in the initial image generation model and the feature image generation model to obtain a second initial object image and a second feature object image; and

fuse the second initial object image and the second feature object image based on the joint mask image to obtain a reference object image, as a result of performing target attribute editing on the second initial object image.

13 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

perform adversarial learning on an initial image generation network and an initial image discrimination network based on the first training image set to obtain an intermediate image generation network and an intermediate image discrimination network;

obtain the initial image generation model based on the intermediate image generation network;

perform adversarial learning on the intermediate image generation network and the intermediate image discrimination network based on the second training image set to obtain a target image generation network; and

obtain the feature image generation model based on the target image generation network.

14 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

acquire a first candidate image set, first candidate images in the first candidate image set being object images corresponding to objects of the category;

perform image alignment on the first candidate images based on positions of a reference object part of the objects of the category in the first candidate images;

obtain the first training image set based on the first candidate images subjected to the image alignment;

acquire a second candidate image set, second candidate images in the second candidate image set being object images corresponding to the objects of the category and comprising the target attribute;

perform image alignment on the second candidate images based on positions of the reference object part of the objects of the category in the second candidate images; and

obtain the second training image set based on the second candidate images subjected to the image alignment.

15 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

use an attribute mask image corresponding to the first initial object image as a first mask image, and use an attribute mask image corresponding to the first feature object image as a second mask image, the first initial object image and the first feature object image having different sizes, and the first mask image and the second mask image having different sizes; and

perform size alignment on the first mask image and the second mask image, and obtain the joint mask image based on the first mask image and the second mask image subjected to the size alignment.

16 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

acquire, from the second initial object image, a first image region matching a shielded region in the joint mask image;

fuse the second initial object image and the second feature object image to obtain a fused object image;

acquire, from the fused object image, a second image region matching a non-shielded region in the joint mask image; and

obtain the reference object image based on the first image region and the second image region.

17 . The apparatus according to claim 16 , wherein the processing circuitry is further configured to

perform size transformation on the joint mask image to obtain a transformed joint mask image of a same size as the second initial object image;

perform re-masking processing on the transformed joint mask image to obtain a re-masked joint mask image;

fuse the second initial object image and the re-masked joint mask image to obtain the first image region; and

fuse the fused object image and the transformed joint mask image to obtain the second image region.

18 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

replace the second initial object image with the reference object image and input the reference object image into downstream network layers of the target network layer in the initial image generation model, and acquire an image output by an end network layer of the initial image generation model as a target object image, wherein

the target object image is a result of performing target attribute editing on an original object image, which is output by the end network layer of the initial image generation model after the value of the first variable is input into the initial image generation model.

19 . The apparatus according to claim 18 , wherein the processing circuitry is further configured to:

replace the second initial object image with the reference object image and input the reference object image into the downstream network layers of the target network layer in the initial image generation model, and acquire an image output by a third network layer in the downstream network layers as a third initial object image;

acquire an image output by a fourth network layer corresponding to the third network layer in the feature image generation model as a third feature object image;

fuse the third initial object image and the third feature object image based on a current joint mask image to obtain an updated object image, the current joint mask image being the joint mask image or an updated joint mask image obtained based on an image output by a fifth network layer in the initial image generation model and an image output by a sixth network layer in the feature image generation model; and

replace the third initial object image with the updated object image and input the updated object image into downstream network layers of the third network layer in the initial image generation model, and acquire an image output by the end network layer of the initial image generation model as the target object image.

20 . A non-transitory computer-readable storage medium storing computer-readable instructions thereon, which, when executed by processing circuitry, cause the processing circuitry to perform an image editing method comprising:

acquiring an initial image generation model and a feature image generation model, the initial image generation model having been trained based on a first training image set, the feature image generation model having been obtained by training the initial image generation model based on a second training image set, object images in the first training image set and the second training image set including objects of a same category, each object image in the second training image set comprising a target attribute, and an image output by the feature image generation model having the target attribute;

based on an image editing request, inputting a value of a first variable as input data into the initial image generation model and the feature image generation model separately;

in response to the input data, acquiring an image output by the initial image generation model to obtain a first initial object image of the category, and acquiring an image output by the feature image generation model to obtain a first feature object image of the category;

based on image regions corresponding to the target attribute in the object images, acquiring attribute mask images respectively corresponding to the first initial object image and the first feature object image, and obtaining a joint mask image based on the attribute mask images;

acquiring images output by corresponding target network layers in the initial image generation model and the feature image generation model to obtain a second initial object image and a second feature object image; and

fusing the second initial object image and the second feature object image based on the joint mask image to obtain a reference object image, as a result of performing target attribute editing on the second initial object image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2025
From: CHEN, HAOKUN; SHEN, RUIXUE; WANG, RUI
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 072735/0161 →
Priority Claims (1)
CN 202211330943.1 · Oct 28, 2022 · national
Continuity (2)
Continuation PCTCN2023119716 · Sep 19, 2023
Related Publication 20240193822A1 · Jun 13, 2024
References Cited (6)
US 20210398334A1 · He · 2021 [cited by examiner]
US 20220207913A1 · Zeng · 2022 [cited by examiner]
US 20230401682A1 · Hu · 2023 [cited by examiner]
CN 111402181A · 2020 [cited by examiner]
CN 115393183A · 2022 [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/CN2023/119716, mailed on Dec. 25, 2023, 14 pages (5 pages of English Translation and 9 pages of Original Document). [cited by applicant]