IP Library Granted Patent US 12,505,520
Granted Patent B2
US 12,505,520 · App. 18/190,654 · Granted Dec 23, 2025

Utilizing a warped digital image with a reposing model to synthesize a modified digital image

Inventors: Krishna Kumar Singh (San Jose, CA); Yijun Li (Seattle, WA); Jingwan Lu (Santa Clara, CA); Duygu Ceylan Aksit (Mountain View, CA); Yangtuanfeng Wang (London, GB); Jimei Yang (Merced, CA); Tobias Hinz (Ulm, DE)
Assignee: Adobe Inc.
G06T5/77G06T3/18G06T7/40G06T7/70G06V10/44G06V10/771G06V10/806G06V10/82G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,520
App. No.
18/190,654
Granted
Dec 23, 2025
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (59)

1 . A computer-implemented method comprising:

generating a texture map from pixel values of a source digital image that depicts a human and from a pose map of the source digital image;

generating a warped digital image by warping the texture map with a target pose map;

combining the warped digital image and the target pose map to generate, utilizing a pose encoder, a warped pose feature map;

generating, utilizing a texture map appearance encoder, a global texture map appearance vector from the texture map; and

synthesizing, utilizing a reposing generative adversarial neural network, a modified digital image that depicts the human according to the target pose map based on the warped pose feature map generated from combining the warped digital image and the target pose map and the global texture map appearance vector generated from the texture map.

2 . The computer-implemented method of claim 1 , wherein generating the texture map further comprises:

generating, utilizing a three-dimensional model, the pose map from the source digital image; and

generating the texture map by projecting the pixel values of the source digital image guided by the pose map of the source digital image.

3 . The computer-implemented method of claim 1 , wherein generating the warped digital image further comprises:

determining UV coordinates from the target pose map; and

generating the warped digital image by rearranging the pixel values from the source digital image in the texture map based on the UV coordinates from the target pose map.

4 . The computer-implemented method of claim 1 , further comprising:

generating, utilizing an image appearance encoder of the reposing generative adversarial neural network, a local appearance vector from the warped digital image;

generating, utilizing a parameter neural network, a local appearance feature tensor from the local appearance vector; and

synthesizing, utilizing the reposing generative adversarial neural network, the modified digital image that depicts the human according to the target pose map based on the warped pose feature map, the global texture map appearance vector, and the local appearance feature tensor.

5 . The computer-implemented method of claim 4 , further comprising generating a globally modified local appearance feature tensor by combining the global texture map appearance vector and the local appearance feature tensor.

6 . The computer-implemented method of claim 5 , further comprising generating, utilizing a style block of the reposing generative adversarial neural network, an intermediate feature vector by modulating the warped pose feature map utilizing the globally modified local appearance feature tensor.

7 . The computer-implemented method of claim 1 , wherein:

generating the warped digital image further comprises generating, utilizing a coordinate inpainting generative neural network, an inpainted warped digital image; and

generating the warped pose feature map comprises generating, utilizing the pose encoder, the warped pose feature map from the inpainted warped digital image.

8 . The computer-implemented method of claim 1 , further comprising training the reposing generative adversarial neural network by:

generating a body mask for the human depicted in the modified digital image; and

determining a measure of loss for a portion of the modified digital image within the body mask utilizing a first loss weight.

9 . The computer-implemented method of claim 8 , further comprising training the reposing generative adversarial neural network by:

determining an additional measure of loss for an additional portion of the modified digital image outside the body mask utilizing a second loss weight; and

modifying parameters of the reposing generative adversarial neural network based on the measure of loss and the additional measure of loss.

10 . A system comprising:

one or more memory devices comprising a source digital image, a target pose map, and a reposing neural network; and

one or more processors configured to cause the system to:

generate a texture map from pixel values of the source digital image that depicts a human and a pose map of the source digital image;

generate a warped digital image by warping the texture map with the target pose map;

combine the warped digital image and the target pose map to generate, utilizing a pose encoder, a warped pose feature map; and

synthesize, utilizing the reposing neural network, a modified digital image that depicts the human according to the target pose map and the warped pose feature map generated from combining the warped digital image and the target pose map.

11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to:

generate, utilizing a three-dimensional model, the pose map from the source digital image;

generate the texture map by projecting the pixel values of the source digital image guided by the pose map of the source digital image; and

generate the warped digital image by rearranging the pixel values from the source digital image in the texture map with UV coordinates from the target pose map.

12 . The system of claim 10 , wherein the one or more processors are configured to cause the system to:

generate, utilizing a coordinate inpainting generative neural network, an inpainted warped digital image from the warped digital image; and

generate, utilizing the pose encoder, the warped pose feature map from the inpainted warped digital image.

13 . The system of claim 10 , wherein the one or more processors are configured to cause the system to:

generate, utilizing a texture map appearance encoder, a global texture map appearance vector from the texture map; and

generate, utilizing an image appearance encoder, a warped image feature vector from the warped digital image.

14 . The system of claim 13 , wherein the one or more processors are configured to cause the system to:

generate, utilizing a parameter neural network, a local appearance feature tensor from the warped image feature vector; and

generate a globally modified local appearance feature tensor by combining the global texture map appearance vector and the local appearance feature tensor.

15 . The system of claim 14 , wherein the one or more processors are configured to cause the system to synthesize the modified digital image that depicts the human according to the target pose map utilizing the reposing neural network from the warped pose feature map and the globally modified local appearance feature tensor.

16 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating a warped digital image by warping a texture map of a digital image with a target pose map;

combining the warped digital image and the target pose map to generate, utilizing a pose encoder, a warped pose feature map;

generating, utilizing a parameter neural network, local appearance feature tensor from the warped digital image; and

synthesizing, utilizing a reposing generative adversarial neural network, a modified digital image that depicts a human according to the target pose map based on the warped pose feature map generated from combining the warped digital image and the target pose map and the local appearance feature tensor.

17 . The non-transitory computer-readable medium of claim 16 , the operations further comprising generating the warped digital image by rearranging pixel values from the digital image in the texture map with UV coordinates from the target pose map.

18 . The non-transitory computer-readable medium of claim 16 , the operations further comprising:

generating a globally modified local appearance feature tensor from the local appearance feature tensor and a global texture map appearance vector from the texture map; and

modulating the warped pose feature map via the reposing generative adversarial neural network utilizing the globally modified local appearance feature tensor.

19 . The non-transitory computer-readable medium of claim 16 , the operations further comprising generating, utilizing a coordinate inpainting generative neural network, an inpainted warped digital image.

20 . The non-transitory computer-readable medium of claim 19 , the operations further comprising generating, utilizing the pose encoder, the warped pose feature map from the inpainted warped digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: SINGH, KRISHNA KUMAR; LI, YIJUN; LU, JINGWAN; AKSIT, DUYGU CEYLAN; WANG, YANGTUANFENG; YANG, JIMEI; HINZ, TOBIAS
To: ADOBE INC.
Reel/Frame 063130/0454 →
Continuity (2)
Provisional Application 63378616 · Oct 6, 2022
Related Publication 20240135513A1 · Apr 25, 2024
References Cited (67)
US 11462040B2 · Lin et al. · 2022 [cited by applicant]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 20190114748A1 · Lin et al. · 2019 [cited by applicant]
US 20200151940A1 · Yu · 2020 [cited by examiner]
US 20200234480A1 · Volkov et al. · 2020 [cited by applicant]
US 20200394828A1 · Shukla et al. · 2020 [cited by applicant]
US 20210056348A1 · Berlin et al. · 2021 [cited by applicant]
US 20210264207A1 · Smith et al. · 2021 [cited by applicant]
US 20210334935A1 · Grigoriev · 2021 [cited by examiner]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20220237829A1 · Ren · 2022 [cited by examiner]
US 20220392133A1 · Volkov et al. · 2022 [cited by applicant]
US 20230037339A1 · Villegas et al. · 2023 [cited by applicant]
US 20230110206A1 · Karras et al. · 2023 [cited by applicant]
US 20230123820A1 · Wang et al. · 2023 [cited by applicant]
US 20230319223A1 · Naruiec et al. · 2023 [cited by applicant]
US 20230410447A1 · Cheng et al. · 2023 [cited by applicant]
US 20240135511A1 · Singh et al. · 2024 [cited by applicant]
US 20240135513A1 · Singh et al. · 2024 [cited by applicant]
US 20240153047A1 · Smith et al. · 2024 [cited by applicant]
US 20240169624A1 · Brandt et al. · 2024 [cited by applicant]
US 20240169701A1 · Kulal et al. · 2024 [cited by applicant]
US 20240171848A1 · Figueroa et al. · 2024 [cited by applicant]
US 20240249459A1 · Bradley et al. · 2024 [cited by applicant]
US 20240331322A1 · Smith · 2024 [cited by applicant]
CN 113240613B · 2021 [cited by applicant]
CN 114862697A · 2022 [cited by applicant]
CN 114943656A · 2022 [cited by applicant]
EP 2238563B1 · 2018 [cited by applicant]
GB 2606253A · 2022 [cited by applicant]
WO 2022083504A1 · 2022 [cited by applicant]
Grigorev et al. “Coordinate-based texture inpainting for pose-guided human image generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. (Year: 2019). [cited by examiner]
Sarkar et al. “Humangan: A generative model of human images.” 2021 International Conference on 3D Vision (3DV). IEEE, 2021. (Year: 2021). [cited by examiner]
Albahar, Badour, et al. “Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN.” arXiv e-prints (2021): arXiv-2109. (Year: 2021). [cited by examiner]
Si et al. “Multistage adversarial losses for pose-based human image synthesis.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. (Year: 2018). [cited by examiner]
Sarkar et al. “Style and pose control for image synthesis of humans from a single monocular view.” arXiv preprint arXiv:2102.11263 (2021). (Year: 2021). [cited by examiner]
Huang et al. “Beyond face rotation: Global and local perception gan for photorealistic and identity preserving frontal view synthesis.” Proceedings of the IEEE international conference on computer vision. 2017. (Year: 2… [cited by examiner]
Xia et al. “Local and global perception generative adversarial network for facial expression synthesis.” IEEE Transactions on Circuits and Systems for Video Technology 32.3 (2021): 1443-1452. (Year: 2021). [cited by examiner]
Liu, Ting, et al. “Spatial-aware texture transformer for high-fidelity garment transfer.” IEEE Transactions on Image Processing 30 (2021): 7499-7510. (Year: 2021). [cited by examiner]
Wang, Tuanfeng Y., et al. “Dance in the wild: Monocular human animation with neural dynamic appearance synthesis.” 2021 International Conference on 3D Vision (3DV). IEEE, 2021. (Year: 2021). [cited by examiner]
Li, Yining, Chen Huang, and Chen Change Loy. “Dense intrinsic appearance flow for human pose transfer.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Dong, Haoye, et al. “Soft-gated warping-gan for pose-guided person image synthesis.” Advances in neural information processing systems 31 (2018). (Year: 2018). [cited by examiner]
Qiao, Fengchun, et al. “Geometry-contrastive gan for facial expression transfer.” arXiv preprint arXiv: 1802.01822 (2018). (Year: 2018). [cited by applicant]
Chen, Yajing, et al. “Self-supervised learning of detailed 3d face reconstruction.” IEEE Transactions on Image Processing 29 (2020): 8696-8705. (Year: 2020). [cited by applicant]
Screen captures from YouTube video clip entitled “How to Use xpression camera—For Video Chat, Vlogging, Live Streaming, Content Creation, Gaming,” 4 pages, uploaded on Feb. 8, 2023 by user “EmbodyMe”. Retrieved from Int… [cited by applicant]
U.S. Appl. No. 18/190,500, Feb. 26, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,500, Apr. 15, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, Mar. 12, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,673, Mar. 13, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, Mar. 12, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, May 7, 2025, Office Action. [cited by applicant]
Michail Christos Doukas, Stefanos Zafeiriou, Viktorija Sharmanska HeadGAN: One-Shot Neural Head Synthesis and Editing. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer—High-Resolution Image Synthesis with Latent Diffusion Models arXiv:2112.10752 Wed, Apr. 13, 2022. [cited by applicant]
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang—Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN Badour et al., Siggraph Asia 2021. [cited by applicant]
Artur Grigorev, Artem Sevastopolsky, Alexander Vakhiov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, and Christian Theobalt. Style and pose control for image synthesis of humans from a single monocular view, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Combined Search and Examination Report as received in GB2318853.5 dated Jun. 12, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319660.3 dated Jun. 14, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319084.6 dated Jun. 25, 2024. [cited by applicant]
Wiles, 0., Koepke, A and Zisserman, A, 2018. “X2face: A network for controlling face generation using images, audio, and pose codes.” in Proceedings of the European conference on computer vision (ECCV) (pp. 690-706). [cited by applicant]
U.S. Appl. No. 18/190,544, Nov. 20, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, Dec. 13, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, Jan. 27, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,636, Mail Date Aug. 26, 2025, Notice of Allowance. [cited by applicant]