IP Library › Granted Patent US 12,333,627
Granted Patent B2
US 12,333,627 · App. 17/718,164 · Granted Jun 17, 2025

Artificial intelligence-based image generation method, device and apparatus, and storage medium

Inventors: Xiaohang Ren (Shenzhen, CN); Yang Ding (Shenzhen, CN); Xiao Zhou (Shenzhen, CN); Youcheng Ben (Shenzhen, CN); Yuxuan Yan (Shenzhen, CN); Pei Cheng (Shenzhen, CN); Gang Yu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T11/00G06T7/73G06V10/40G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,627
App. No.
17/718,164
Granted
Jun 17, 2025
Kind
B2
Abstract

An embodiment of this application discloses an artificial intelligence-based image generation method performed by a computer device. The method includes: acquiring a source image including a target object whose pose is to be transformed, and a target image including a reference object presenting a target pose; determining a pose transition matrix according to a model pose corresponding to the pose of the target object and a model pose corresponding to the target pose of the reference object; extracting a basic appearance feature of the target object from the source image; processing the basic appearance feature based on the pose transition matrix, to obtain a target appearance feature of the target object in the target pose; and generating a target synthetic image of the target object in the target pose based on the target appearance feature.

Claims (73)

1. An artificial intelligence-based image generation method performed by a computer device, the method comprising:

acquiring a 2D source image and a 2D target image, the 2D source image comprising a target object whose pose is to be transformed, and the 2D target image comprising a reference object presenting a target pose, wherein the reference object is different from the target object;

determining a 3D pose transition matrix according to a 3D model pose corresponding to the to-be-transformed pose of the target object and a 3D model pose corresponding to the target pose of the reference object;

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 2D source image;

processing the 3D basic appearance feature based on the 3D pose transition matrix, to obtain a 3D target appearance feature of the target object in the target pose;

generating a 2D target synthetic image of the target object in the target pose based on the 3D target appearance feature; and

outputting the 2D target synthetic image on a display of the computer device.

2. The method according to claim 1 , wherein the extracting comprises:

extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image; and

the generating comprises:

generating a 2D target synthetic image of the target object in the target pose by the generator based on the 3D target appearance feature.

3. The method according to claim 2 , wherein the extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image comprises:

determining a 3D global feature of the source image by the generator; and

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 3D global feature of the 2D source image.

4. The method according to claim 3 , further comprising:

determining positions of appearance feature distribution areas respectively corresponding to N target appearance feature sites on the target object in the 2D source image, wherein N is an integer greater than 1; and

wherein the extracting, as the 3D basic appearance feature, an appearance feature of the target object from the 3D global feature of the 2D source image comprises:

extracting, according to the positions of appearance feature distribution areas respectively corresponding to the N target appearance feature sites, local features respectively corresponding to the N target appearance feature sites from the 3D global feature of the 2D source image, to form the 3D basic appearance feature.

5. The method according to claim 2 , wherein the generating comprises:

acquiring a pose feature of the reference object in the 2D target image, wherein the pose feature of the reference object is extracted from a 3D global feature of the target image, and the 3D global feature of the 2D target image is determined by the generator; and

generating a 2D target synthetic image of the target object in the target pose by the generator based on the pose feature of the reference object and the 3D target appearance feature.

6. The method according to claim 1 , wherein the 2D target image is one of a plurality of target video frames in a target action video, and the 2D target synthetic image corresponds to each target video frame; and

after generating a 2D target synthetic image of the target object in a target pose respectively corresponding to each target video frame in the target action video, the method further comprises:

arranging the 2D target synthetic images respectively corresponding to the target video frames according to the sequence of the target video frames in the target action video to obtain a target synthetic video of the 2D target synthetic images, each target synthetic image including the target object in a target pose corresponding to a corresponding one of the target video frames.

7. The method according to claim 1 , wherein the 3D model pose corresponding to the pose of the target object comprises a 3D model corresponding to the target object having the pose to be transformed; and the 3D model pose corresponding to the target pose of the reference object comprises a 3D model corresponding to the reference object having the target pose.

8. A computer device comprising a processor and a storage,

the storage being configured to store a plurality of computer programs; and

the processor being configured to execute the plurality of computer programs to implement an artificial intelligence-based image generation method including:

acquiring a 2D source image and a 2D target image, the 2D source image comprising a target object whose pose is to be transformed, and the 2D target image comprising a reference object presenting a target pose, wherein the reference object is different from the target object;

determining a 3D pose transition matrix according to a 3D model pose corresponding to the to-be-transformed pose of the target object and a 3D model pose corresponding to the target pose of the reference object;

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 2D source image;

processing the 3D basic appearance feature based on the 3D pose transition matrix, to obtain a 3D target appearance feature of the target object in the target pose;

generating a 2D target synthetic image of the target object in the target pose based on the 3D target appearance feature; and

outputting the 2D target synthetic image on a display of the computer device.

9. The computer device according to claim 8 , wherein the extracting comprises:

extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image; and

the generating comprises:

generating a 2D target synthetic image of the target object in the target pose by the generator based on the 3D target appearance feature.

10. The computer device according to claim 9 , wherein the extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image comprises:

determining a 3D global feature of the source image by the generator; and

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 3D global feature of the 2D source image.

11. The computer device according to claim 10 , wherein the method further comprises:

determining positions of appearance feature distribution areas respectively corresponding to N target appearance feature sites on the target object in the 2D source image, wherein N is an integer greater than 1; and

wherein the extracting, as the 3D basic appearance feature, an appearance feature of the target object from the 3D global feature of the 2D source image comprises:

extracting, according to the positions of appearance feature distribution areas respectively corresponding to the N target appearance feature sites, local features respectively corresponding to the N target appearance feature sites from the 3D global feature of the 2D source image, to form the 3D basic appearance feature.

12. The computer device according to claim 9 , wherein the generating comprises:

acquiring a pose feature of the reference object in the 2D target image, wherein the pose feature of the reference object is extracted from a 3D global feature of the target image, and the 3D global feature of the 2D target image is determined by the generator; and

generating a 2D target synthetic image of the target object in the target pose by the generator based on the pose feature of the reference object and the 3D target appearance feature.

13. The computer device according to claim 8 , wherein the target image is one of a plurality of target video frames in a target action video, and the target synthetic image corresponds to each target video frame; and

after generating a 2D target synthetic image of the target object in a target pose respectively corresponding to each target video frame in the target action video, the method further comprises:

arranging the 2D target synthetic images respectively corresponding to the target video frames according to the sequence of the target video frames in the target action video to obtain a target synthetic video of the 2D target synthetic images, each target synthetic image including the target object in a target pose corresponding to a corresponding one of the target video frames.

14. The computer device according to claim 8 , wherein the 3D model pose corresponding to the pose of the target object comprises a 3D model corresponding to the target object having the pose to be transformed; and the 3D model pose corresponding to the target pose of the reference object comprises a 3D model corresponding to the reference object having the target pose.

15. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium being configured to store a plurality of computer programs, and the plurality of computer programs, when executed by a processor of a computer device, causing the computer device to implement an artificial intelligence-based image generation method including:

acquiring a 2D source image and a 2D target image, the 2D source image comprising a target object whose pose is to be transformed, and the 2D target image comprising a reference object presenting a target pose, wherein the reference object is different from the target object;

determining a 3D pose transition matrix according to a 3D model pose corresponding to the to-be-transformed pose of the target object and a 3D model pose corresponding to the target pose of the reference object;

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 2D source image;

processing the 3D basic appearance feature based on the 3D pose transition matrix, to obtain a 3D target appearance feature of the target object in the target pose;

generating a 2D target synthetic image of the target object in the target pose based on the 3D target appearance feature; and

outputting the 2D target synthetic image on a display of the computer device.

16. The non-transitory computer-readable storage medium according to claim 15 , wherein the extracting comprises:

extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image; and

the generating comprises:

generating a 2D target synthetic image of the target object in the target pose by the generator based on the 3D target appearance feature.

17. The non-transitory computer-readable storage medium according to claim 16 , wherein the extracting, as a 3D basic appearance feature, an appearance feature of the target object by a generator from the 2D source image comprises:

determining a 3D global feature of the source image by the generator; and

extracting, as a 3D basic appearance feature, an appearance feature of the target object from the 3D global feature of the 2D source image.

18. The non-transitory computer-readable storage medium according to claim 16 , wherein the generating comprises:

acquiring a pose feature of the reference object in the 2D target image, wherein the pose feature of the reference object is extracted from a 3D global feature of the target image, and the 3D global feature of the 2D target image is determined by the generator; and

generating a 2D target synthetic image of the target object in the target pose by the generator based on the pose feature of the reference object and the 3D target appearance feature.

19. The non-transitory computer-readable storage medium according to claim 15 , wherein the target image is one of a plurality of target video frames in a target action video, and the target synthetic image corresponds to each target video frame; and

after generating a 2D target synthetic image of the target object in a target pose respectively corresponding to each target video frame in the target action video, the method further comprises:

arranging the 2D target synthetic images respectively corresponding to the target video frames according to the sequence of the target video frames in the target action video to obtain a target synthetic video of the 2D target synthetic images, each target synthetic image including the target object in a target pose corresponding to a corresponding one of the target video frames.

20. The non-transitory computer-readable storage medium according to claim 15 , wherein the 3D model pose corresponding to the pose of the target object comprises a 3D model corresponding to the target object having the pose to be transformed; and the 3D model pose corresponding to the target pose of the reference object comprises a 3D model corresponding to the reference object having the target pose.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2022
From: YU, GANG; ZHOU, XIAO; CHENG, PEI; DING, YANG; REN, XIAOHANG; BEN, YOUCHENG; YAN, YUXUAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 060795/0703 →
Priority Claims (1)
CN 202010467388.1 · May 28, 2020 · national
Continuity (2)
Continuation PCTCN2021091825 · May 6, 2021
Related Publication 20220237829A1 · Jul 28, 2022
References Cited (18)
US 10135835B1 · Kandel et al. · 2018 [cited by applicant]
US 10250395B1 · Borne-Pons et al. · 2019 [cited by applicant]
US 11024060B1 · Ma · 2021 [cited by examiner]
US 20190228556A1 · Wang · 2019 [cited by examiner]
CN 102137111A · 2011 [cited by applicant]
CN 105488665A · 2016 [cited by applicant]
CN 106060036A · 2016 [cited by applicant]
CN 106484511A · 2017 [cited by applicant]
CN 107465693A · 2017 [cited by applicant]
CN 108564119A · 2018 [cited by applicant]
CN 110930578A · 2020 [cited by applicant]
CN 111027438A · 2020 [cited by applicant]
CN 111047548A · 2020 [cited by applicant]
CN 111626218A · 2020 [cited by applicant]
Tencent Technology, WO, PCT/CN2021/091825, Aug. 5, 2021, 6 pgs. [cited by applicant]
Tencent Technology, IPRP, PCT/CN2021/091825, Nov. 17, 2022, 7 pgs. [cited by applicant]
Tencent Technology, Indian Office Action, IN Patent Application No. 202247037535, Feb. 2, 2023, 7 pgs. [cited by applicant]
Tencent Technology, ISR, PCT/CN2021/091825, Aug. 5, 2021, 2 pgs. [cited by applicant]