IP Library › Granted Patent US 12,456,169
Granted Patent B2
US 12,456,169 · App. 18/129,213 · Granted Oct 28, 2025

Image processing method and apparatus

Inventors: Feida Zhu (Beijing, CN); Ying Tai (Beijing, CN); Chengjie Wang (Beijing, CN); Jilin Li (Beijing, CN)
Assignee: Tencent Cloud Computing (Beijing) Co., Ltd.
G06T5/50G06T5/77G06T7/11G06T9/00G06V10/25G06V10/761G06T2207/20081G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,169
App. No.
18/129,213
Granted
Oct 28, 2025
Kind
B2
Abstract

This disclosure relates to image processing method and apparatus. The method includes: processing, with a first generator in an image processing model, a first sample image in a first sample set to obtain a first predicted image; processing, with the first generator, a second sample image in a second sample set to obtain a second predicted image; and training the image processing model according to a difference between the target avatar in the first sample image and the first predicted avatar and a difference between a first type attribute of the sample avatar in the second sample image and a first type attribute of the second predicted avatar.

Claims (61)

1. An image processing method, comprising:

processing, with a first generator in an image processing model, a first sample image x i in a first sample set to obtain a first predicted image x′ i , the first predicted image x′ i comprising a first predicted avatar, the first sample set comprising N first sample images, each of the first sample images comprising a target avatar corresponding to a same target user, N being a positive integer, i being a positive integer, and i≤N;

processing, with the first generator, a second sample image y k in a second sample set to obtain a second predicted image y′ k , the second predicted image y′ k comprising a second predicted avatar; the second sample set comprising M second sample images, each of the second sample images comprising a sample avatar, M being a positive integer, k being a positive integer, and k≤M; and

training the image processing model according to a difference between the target avatar in the first sample image x i and the first predicted avatar and a difference between a first type attribute of the sample avatar in the second sample image y k and a first type attribute of the second predicted avatar, the image processing model being configured to replace an avatar in an input image with the target avatar and retain the first type attribute of the avatar in the input image.

2. The method according to claim 1 , wherein the first generator comprises an encoder and a first decoder, and the processing the first sample image x i in the first sample set to obtain the first predicted image x′ i comprises:

encoding, with the encoder, the first sample image x i to obtain a first feature vector; and

decoding, with the first decoder, the first feature vector to obtain a first generated image and first region segmentation information, the first region segmentation information being for indicating an avatar region in the first generated image.

3. The method according to claim 2 , the processing the first sample image x i in the first sample set to obtain the first predicted image x′ i comprises:

extracting the first predicted image x′ i from the first generated image according to the first region segmentation information.

4. The method according to claim 1 , wherein the first generator comprises an encoder and a first decoder, and the processing the second sample image y k in the second sample set to obtain the second predicted image y′ k comprises:

encoding, with the encoder, the second sample image y k to obtain a second feature vector; and

decoding, with the first decoder, the second feature vector to obtain a second generated image and second region segmentation information, the second region segmentation information being for indicating an avatar region in the second generated image.

5. The method according to claim 4 , wherein the processing the second sample image y k in the second sample set to obtain the second predicted image y′ k comprises:

extracting the second predicted image y′ k from the second generated image according to the second region segmentation information.

6. The method according to claim 1 , further comprising:

acquiring a test video, the test video comprising R frames of test images, each frame of test image comprising a corrected avatar, and R being a positive integer;

processing, with the first generator of the trained image processing model, the R frames of test images to obtain R frames of predicted images respectively corresponding to the R frames of test images, wherein the R frames of predicted images comprise the target avatar of the target user, and the first type attribute of the avatars of the R frames of predicted images is kept consistent with the first type attribute of the corrected avatar in the corresponding test image; and

performing image inpainting on the R frames of test images from which the corrected avatars are deleted in the test video.

7. The method according to claim 6 , further comprising:

after the image inpainting, fusing the R frames of test images with the corresponding test images in the test video respectively to obtain a target video.

8. The method according to claim 1 , wherein the first type attribute comprises a non-identity recognition attribute.

9. An image processing apparatus, comprising:

a memory operable to store computer-readable instructions; and

a processor circuitry operable to read the computer-readable instructions, the processor circuitry when executing the computer-readable instructions is configured to:

process, with a first generator in an image processing model, a first sample image x i in a first sample set to obtain a first predicted image x′ i , the first predicted image x′ i comprising a first predicted avatar, the first sample set comprising N first sample images, each of the first sample images comprising a target avatar corresponding to a same target user, N being a positive integer, i being a positive integer, and i≤N;

process, with the first generator, a second sample image y k in a second sample set to obtain a second predicted image y′ k , the second predicted image y′ k comprising a second predicted avatar; the second sample set comprising M second sample images, each of the second sample images comprising a sample avatar, M being a positive integer, k being a positive integer, and k≤M; and

train the image processing model according to a difference between the target avatar in the first sample image x i and the first predicted avatar and a difference between a first type attribute of the sample avatar in the second sample image y k and a first type attribute of the second predicted avatar, the image processing model being configured to replace an avatar in an input image with the target avatar and retain the first type attribute of the avatar in the input image.

10. The apparatus according to claim 9 , wherein the first generator comprises an encoder and a first decoder, and the processor circuitry is configured to:

encode, with the encoder, the first sample image x i to obtain a first feature vector; and

decode, with the first decoder, the first feature vector to obtain a first generated image and first region segmentation information, the first region segmentation information being for indicating an avatar region in the first generated image.

11. The apparatus according to claim 10 , wherein the processor circuitry is configured to:

extract the first predicted image x′ i from the first generated image according to the first region segmentation information.

12. The apparatus according to claim 9 , wherein the first generator comprises an encoder and a first decoder, and the processor circuitry is configured to:

encode, with the encoder, the second sample image y k to obtain a second feature vector;

decode, with the first decoder, the second feature vector to obtain a second generated image and second region segmentation information, the second region segmentation information being for indicating an avatar region in the second generated image.

13. The apparatus according to claim 12 , wherein the processor circuitry is configured to:

extract the second predicted image y′ k from the second generated image according to the second region segmentation information.

14. The apparatus according to claim 9 , the processor circuitry is further configured to:

acquire a test video, the test video comprising R frames of test images, each frame of test image comprising a corrected avatar, and R being a positive integer;

process, with the first generator of the trained image processing model, the R frames of test images to obtain R frames of predicted images respectively corresponding to the R frames of test images, wherein the R frames of predicted images comprise the target avatar of the target user, and the first type attribute of the avatars of the R frames of predicted images is kept consistent with the first type attribute of the corrected avatar in the corresponding test image; and

perform image inpainting on the R frames of test images from which the corrected avatars are deleted in the test video.

15. The apparatus according to claim 14 , the processor circuitry is further configured to:

after the image inpainting, fuse the R frames of test images with the corresponding test images in the test video respectively to obtain a target video.

16. A non-transitory machine-readable media, having instructions stored on the machine-readable media, the instructions configured to, when executed, cause a machine to:

process, with a first generator in an image processing model, a first sample image x i in a first sample set to obtain a first predicted image x′ i , the first predicted image x′ i comprising a first predicted avatar, the first sample set comprising N first sample images, each of the first sample images comprising a target avatar corresponding to a same target user, N being a positive integer, i being a positive integer, and i≤N;

process, with the first generator, a second sample image y k in a second sample set to obtain a second predicted image y k ′, the second predicted image y′ k comprising a second predicted avatar; the second sample set comprising M second sample images, each of the second sample images comprising a sample avatar, M being a positive integer, k being a positive integer, and k≤M; and

train the image processing model according to a difference between the target avatar in the first sample image x i and the first predicted avatar and a difference between a first type attribute of the sample avatar in the second sample image y k and a first type attribute of the second predicted avatar, the image processing model being configured to replace an avatar in an input image with the target avatar and retain the first type attribute of the avatar in the input image.

17. The non-transitory machine-readable media according to claim 16 , wherein the first generator comprises an encoder and a first decoder, and the instructions are configured to cause the machine to:

encode, with the encoder, the first sample image x i to obtain a first feature vector;

decode, with the first decoder, the first feature vector to obtain a first generated image and first region segmentation information, the first region segmentation information being for indicating an avatar region in the first generated image; and

extract the first predicted image x′ i from the first generated image according to the first region segmentation information.

18. The non-transitory machine-readable media according to claim 16 , wherein the first generator comprises an encoder and a first decoder, and the instructions are configured to cause the machine to:

encode, with the encoder, the second sample image y k to obtain a second feature vector;

decode, with the first decoder, the second feature vector to obtain a second generated image and second region segmentation information, the second region segmentation information being for indicating an avatar region in the second generated image; and

extract the second predicted image y′ k from the second generated image according to the second region segmentation information.

19. The non-transitory machine-readable media according to claim 16 , the instructions are further configured to cause the machine to:

acquire a test video, the test video comprising R frames of test images, each frame of test image comprising a corrected avatar, and R being a positive integer;

process, with the first generator of the trained image processing model, the R frames of test images to obtain R frames of predicted images respectively corresponding to the R frames of test images, wherein the R frames of predicted images comprise the target avatar of the target user, and the first type attribute of the avatars of the R frames of predicted images is kept consistent with the first type attribute of the corrected avatar in the corresponding test image; and

perform image inpainting on the R frames of test images from which the corrected avatars are deleted in the test video.

20. The non-transitory machine-readable media according to claim 19 , the instructions are further configured to cause the machine to:

after the image inpainting, fuse the R frames of test images with the corresponding test images in the test video respectively to obtain a target video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: ZHU, FEIDA; TAI, YING; WANG, CHENGJIE; LI, JILIN
To: TENCENT CLOUD COMPUTING (BEIJING) CO., LTD.
Reel/Frame 063193/0362 →
Priority Claims (1)
CN 2021106203828 · Jun 3, 2021 · national
Continuity (2)
Continuation PCTCN2021108489 · Jul 26, 2021
Related Publication 20230237630A1 · Jul 27, 2023
References Cited (20)
US 20200349391A1 · Zhang · 2020 [cited by examiner]
US 20200372351A1 · Chang et al. · 2020 [cited by applicant]
CN 110084193A · 2019 [cited by applicant]
CN 111402118 · 2020 [cited by applicant]
CN 111523413 · 2020 [cited by applicant]
CN 111523413A · 2020 [cited by applicant]
CN 111783603 · 2020 [cited by applicant]
CN 111860362 · 2020 [cited by applicant]
CN 112017140A · 2020 [cited by applicant]
JP 202143839A · 2021 [cited by applicant]
Office action issued in Korean application No. 10-2023-7030383, dated Oct. 16, 2024, 6 pages (with English translation). [cited by applicant]
Office action issued in Japanese applicatin No. 2023-576240, dated Dec. 3, 2024, 10 pages (with English translation). [cited by applicant]
Yang et al., “Face Swapping Neural Networks Based on Improved Autoencoders,” IEEE 5th International Conference on Big Data Intelligence and Computing, (DATACOM), Nov. 2019, 107-112. [cited by applicant]
Yan et al., “Video Face Swap Based on Autoencoder Generation Network,” International Conference on Audio, Language and Image Processing (ICALIP), Jul. 2018, 103-108. [cited by applicant]
Wiles et al., “X2Face: A Network for Controlling Face Generation Using Images, Audio, and Pose Codes,” In Proceedings of the European Conference on Computer Vision, Oct. 2018, pp. 690-706. [cited by applicant]
Li et al., “Advancing High Fidelity Identity Swapping for Forgery Detection,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 5074-5083. [cited by applicant]
Burkov et al., “Neural Head Reenactment with Latent Pose Descriptors,” In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 2020, pp. 13783-13792. [cited by applicant]
Extended European Search Report in European application No. 21943728.2, dated Jul. 23, 2024, 6 pages. [cited by applicant]
Office action in Japanese application No. 2023-576240, dated Jul. 16, 2024, 10 pages (with English translation). [cited by applicant]
International Search Report issued Feb. 24, 2022 in International (PCT) Application No. PCT/CN2021/108489. [cited by applicant]