IP Library Granted Patent US 12,462,420
Granted Patent B2
US 12,462,420 · App. 18/190,636 · Granted Nov 4, 2025

Synthesizing a modified digital image utilizing a reposing model

Inventors: Krishna Kumar Singh (San Jose, CA); Yijun Li (Seattle, WA); Jingwan Lu (Santa Clara, CA); Duygu Ceylan Aksit (Mountain View, CA); Yangtuanfeng Wang (London, GB); Jimei Yang (Merced, CA); Tobias Hinz (Ulm, DE)
Assignee: Adobe Inc.
G06T7/70G06T7/40G06V10/44G06V10/771G06V10/806G06V10/82G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,420
App. No.
18/190,636
Filed
Mar 27, 2023
Granted
Nov 4, 2025
Kind
B2
Art Unit
2664
USPC
382/103
Abstract

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.

Claims (58)

1 . A computer-implemented method comprising:

generating, utilizing a pose encoder, a pose feature map from a target pose map and a source digital image that depicts a human;

generating, utilizing a texture map appearance encoder, a global texture map appearance vector from a texture map of the source digital image;

generating, utilizing a parameter neural network, a local appearance feature tensor from the source digital image; and

synthesizing, utilizing a reposing generative adversarial neural network, a modified digital image that depicts the human according to the target pose map based on the pose feature map, the local appearance feature tensor, and the global texture map appearance vector.

2 . The computer-implemented method of claim 1 , wherein generating the pose feature map, utilizing the pose encoder comprises generating the pose feature map utilizing a hierarchical pose encoder by:

generating, utilizing a downsampling layer of the hierarchical pose encoder, an intermediate downsampled feature vector having a resolution;

generating utilizing an upsampling layer of the hierarchical pose encoder, an intermediate upsampled feature vector having the resolution; and

generating, via a skip connection, a combined feature vector from the intermediate downsampled feature vector and the intermediate upsampled feature vector.

3 . The computer-implemented method of claim 2 , wherein generating the pose feature map utilizing the hierarchical pose encoder further comprises:

generating, utilizing an additional downsampling layer of the hierarchical pose encoder, an additional intermediate downsampled feature vector having an additional resolution;

generating utilizing an additional upsampling layer of the hierarchical pose encoder from the combined feature vector, an additional intermediate upsampled feature vector having the additional resolution; and

combining, via a skip connection, the additional intermediate downsampled feature vector and the additional intermediate upsampled feature vector to generate the pose feature map.

4 . The computer-implemented method of claim 1 , wherein generating the global texture map appearance vector from the texture map of the source digital image comprises generating a one-dimensional appearance vector utilizing a hierarchical appearance encoder.

5 . The computer-implemented method of claim 1 , wherein generating the local appearance feature tensor further comprises:

generating, utilizing an image appearance encoder, a local appearance vector from the source digital image; and

generating, utilizing the parameter neural network, the local appearance feature tensor from the local appearance vector, the local appearance feature tensor comprising a spatially varying scaling tensor and a spatially varying shifting tensor.

6 . The computer-implemented method of claim 1 , wherein synthesizing the modified digital image utilizing the reposing generative adversarial neural network comprises generating the modified digital image from the pose feature map by:

generating globally modified local appearance feature tensor by combining the local appearance feature tensor and the global texture map appearance vector; and

modulating the reposing generative adversarial neural network utilizing the globally modified local appearance feature tensor.

7 . The computer-implemented method of claim 1 , further comprising training the reposing generative adversarial neural network by:

generating incomplete digital images from an unpaired image dataset by applying masks to a plurality of source digital images portraying humans in the unpaired image dataset; and

synthesizing, utilizing the reposing generative adversarial neural network, a plurality of inpainted digital images from the incomplete digital images and target pose maps of the plurality of source digital images.

8 . The computer-implemented method of claim 7 , further comprising modifying parameters of the reposing generative adversarial neural network based on measures of loss between the plurality of inpainted digital images and the plurality of source digital images.

9 . The computer-implemented method of claim 1 , further comprising training the reposing generative adversarial neural network utilizing an unpaired dataset comprising the source digital image by:

determining, utilizing a discriminator neural network, an adversarial loss from the modified digital image; and

modifying parameters of the reposing generative adversarial neural network based on the adversarial loss from the modified digital image generated from the unpaired dataset.

10 . A system comprising:

one or more memory devices comprising a source digital image that depicts a human, a texture map of the source digital image, a target pose map of the human, and a reposing neural network; and

one or more processors configured to cause the system to:

generate, utilizing a texture map appearance encoder of the reposing neural network, a global texture map appearance vector from the texture map;

generate, utilizing an image appearance encoder of the reposing neural network, a local appearance vector from the source digital image;

generate, utilizing a parameter neural network, local appearance feature tensor from the local appearance vector; and

synthesize, utilizing the reposing neural network, a modified digital image that depicts the human according to the target pose map based on the source digital image, the local appearance feature tensor, and the global texture map appearance vector.

11 . The system of claim 10 , wherein the one or more processors are configured to cause the system to:

generate the global texture map appearance vector from the texture map of the source digital image comprises generating a one-dimensional appearance vector utilizing a hierarchical appearance encoder; and

generate the local appearance feature tensor from the local appearance vector comprises a spatially varying scaling tensor and a spatially varying shifting tensor.

12 . The system of claim 10 , wherein the one or more processors are configured to cause the system to generate globally modified local appearance feature tensor by combining the global texture map appearance vector with the local appearance feature tensor.

13 . The system of claim 12 , wherein the one or more processors are configured to cause the system to modulate the reposing neural network utilizing the globally modified local appearance feature tensor.

14 . The system of claim 10 , wherein the one or more processors are configured to cause the system to generate the texture map from pixel values of the source digital image and a pose map of the source digital image.

15 . The system of claim 10 , wherein the one or more processors are configured to cause the system to:

train the reposing neural network by:

generating incomplete digital images from an unpaired image dataset by applying masks to a plurality of source digital images portraying humans in the unpaired image dataset;

synthesizing, utilizing the reposing neural network, a plurality of inpainted digital images from the incomplete digital images and target pose maps of the plurality of source digital images; and

modifying parameters of the reposing neural network based on measures of loss between the plurality of inpainted digital images and the plurality of source digital images.

16 . A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

generating, utilizing a texture map appearance encoder of a reposing neural network, a global texture map appearance vector from a texture map;

generating, utilizing an image appearance encoder of the reposing neural network, a local appearance vector from a source digital image that depicts a human;

generating, a combined vector from the global texture map appearance vector and the local appearance vector; and

synthesizing, utilizing a reposing generative adversarial neural network, a modified digital image that depicts the human according to a target pose map based on the combined vector.

17 . The non-transitory computer-readable medium of claim 16 , the operations further comprising:

generating the texture map from pixel values of the source digital image and a pose map of the source digital image; and

generating the global texture map appearance vector from the texture map of the source digital image comprises generating a one-dimensional appearance vector utilizing a hierarchical appearance encoder.

18 . The non-transitory computer-readable medium of claim 16 , the operations further comprising:

generating a local appearance feature tensor from the local appearance vector, the local appearance feature tensor comprises a spatially varying scaling tensor and a spatially varying shifting tensor; and

generating the combined vector by combining the global texture map appearance vector with the local appearance feature tensor.

19 . The non-transitory computer-readable medium of claim 16 , the operations further comprising synthesizing the modified digital image by modulating a style block of the reposing neural network utilizing the combined vector.

20 . The non-transitory computer-readable medium of claim 16 , the operations further comprising training the reposing neural network by modifying parameters of the reposing neural network based on an adversarial loss from the modified digital image generated from an unpaired dataset comprising the source digital image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: SINGH, KRISHNA KUMAR; LI, YIJUN; LU, JINGWAN; AKSIT, DUYGU CEYLAN; WANG, YANGTUANFENG; YANG, JIMEI; HINZ, TOBIAS
To: ADOBE INC.
Reel/Frame 063124/0680 →
Continuity (2)
Provisional Application 63378616 · Oct 6, 2022
Related Publication 20240135572A1 · Apr 25, 2024
References Cited (61)
US 11462040B2 · Lin et al. · 2022 [cited by applicant]
US 12026845B2 · Pardeshi · 2024 [cited by applicant]
US 20190114748A1 · Lin et al. · 2019 [cited by applicant]
US 20200234480A1 · Volkov et al. · 2020 [cited by applicant]
US 20200394828A1 · Shukla et al. · 2020 [cited by applicant]
US 20210056348A1 · Berlin et al. · 2021 [cited by applicant]
US 20210264207A1 · Smith et al. · 2021 [cited by applicant]
US 20210334935A1 · Grigoriev et al. · 2021 [cited by applicant]
US 20220068037A1 · Pardeshi · 2022 [cited by applicant]
US 20220207262A1 · Jeong et al. · 2022 [cited by applicant]
US 20220237829A1 · Ren et al. · 2022 [cited by applicant]
US 20220392133A1 · Volkov et al. · 2022 [cited by applicant]
US 20230037339A1 · Villegas et al. · 2023 [cited by applicant]
US 20230110206A1 · Karras et al. · 2023 [cited by applicant]
US 20230123820A1 · Wang et al. · 2023 [cited by applicant]
US 20230319223A1 · Naruiec et al. · 2023 [cited by applicant]
US 20230410447A1 · Cheng et al. · 2023 [cited by applicant]
US 20240135511A1 · Singh et al. · 2024 [cited by applicant]
US 20240135513A1 · Singh et al. · 2024 [cited by applicant]
US 20240153047A1 · Smith et al. · 2024 [cited by applicant]
US 20240169624A1 · Brandt et al. · 2024 [cited by applicant]
US 20240169701A1 · Kulal et al. · 2024 [cited by applicant]
US 20240171848A1 · Figueroa et al. · 2024 [cited by applicant]
US 20240249459A1 · Bradley et al. · 2024 [cited by applicant]
US 20240331322A1 · Smith · 2024 [cited by applicant]
CN 113240613B · 2021 [cited by applicant]
CN 114862697A · 2022 [cited by applicant]
CN 114943656A · 2022 [cited by applicant]
EP 2238563B1 · 2018 [cited by examiner]
GB 2606253A · 2022 [cited by applicant]
WO 2022083504A1 · 2022 [cited by applicant]
Michail Christos Doukas, Stefanos Zafeiriou, Viktorija Sharmanska HeadGAN: One-Shot Neural Head Synthesis and Editing. [cited by applicant]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer—High-Resolution Image Synthesis with Latent Diffusion Models arXiv:2112.10752 Wed, Apr. 13, 2022. [cited by applicant]
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang—Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN Badour et al., SIGGRAPH Asia 2021. [cited by applicant]
Artur Grigorev, Artem Sevastopolsky, Alexander Vakhiov, and Victor Lempitsky. Coordinate-based texture inpainting for pose-guided image generation. arXiv preprint arXiv:1811.11459, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Kripasindhu Sarkar, Vladislav Golyanik, Lingjie Liu, and Christian Theobalt. Style and pose control for image synthesis of humans from a single monocular view, 2021. [cited by applicant]
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [cited by applicant]
Combined Search and Examination Report as received in GB2318853.5 dated Jun. 12, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319660.3 dated Jun. 14, 2024. [cited by applicant]
Combined Search and Examination Report as received in GB2319084.6 dated Jun. 25, 2024. [cited by applicant]
Wiles, 0., Koepke, A and Zisserman, A, 2018. “X2face: A network for controlling face generation using images, audio, and pose codes.” in Proceedings of the European conference on computer vision (ECCV) (pp. 690-706). [cited by applicant]
U.S. Appl. No. 18/190,544, filed Nov. 20, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Dec. 13, 2024, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, filed Jan. 27, 2025, Office Action. [cited by applicant]
Qiao, Fengchun, et al. “Geometry-contrastive gan for facial expression transfer.” arXiv preprint arXiv: 1802.01822 (2018). (Year: 2018). [cited by applicant]
Chen, Yajing, et al. “Self-supervised learning of detailed 3d face reconstruction.” IEEE Transactions on Image Processing 29 (2020): 8696-8705. (Year: 2020). [cited by applicant]
Screen captures from YouTube video clip entitled “How to Use xpression camera—For Video Chat, Vlogging, Live Streaming, Content Creation, Gaming,” 4 pages, uploaded on Feb. 8, 2023 by user “EmbodyMe”. Retrieved from Int… [cited by applicant]
U.S. Appl. No. 18/190,500, field Feb. 26, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,500, filed Apr. 15, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,556, filed Mar. 12, 2025, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 18/190,673, filed Mar. 13, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, filed Mar. 12, 2025, Office Action. [cited by applicant]
U.S. Appl. No. 18/190,684, filed May 7, 2025, Office Action. [cited by applicant]
Albahar, Badour, et al. “Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN.” arXiv e-prints (2021): arXiv-2109. (Year: 2021). [cited by applicant]
Grigorev, Artur et al. “Coordinate-based texture inpainting for pose-guided human image generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. (Year: 2019). [cited by applicant]
Huang, Rui et al. “Beyond face rotation: Global and local perception GAN for photorealistic and identity preserving frontal view synthesis. ”Proceedings of the IEEE international conference on computer vision. 2017. (Ye… [cited by applicant]
Sarkar, Kripasindhu et al. “HumanGAN: A generative model of human images.” 2021 International Conference on 3D Vision (3DV). IEEE, 2021. (Year: 2021). [cited by applicant]
Sarkar, Kripasindhu et al. “Style and pose control for image synthesis of humans from a single monocular view.” arXiv preprint arXiv:2102.11263 (2021). (Year: 2021). [cited by applicant]
Si, Chenyang et al. “Multistage adversarial losses for pose-based human image synthesis.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. (Year: 2018). [cited by applicant]
Xia, Yifan et al. “Local and global perception generative adversarial network for facial expression synthesis.” IEEE Transactions on Circuits and Systems for Video Technology 32.3 (2021): 1443-1452. (Year: 2021). [cited by applicant]