IP Library Granted Patent US 12,675,917
Granted Patent B2
US 12,675,917 · App. 18/320,857 · Granted Jul 7, 2026

Identity-preserving image generation using diffusion models

Inventors: Manuel Jakob Kansy (Zurich, CH); Anton Julien Raël (Zurich, CH); Jacek Krzysztof Naruniec (Windlach, CH); Christopher Richard Schroers (Uster, CH); Romann Matthew Weber (Uster, CH)
Assignees: Disney Enterprises, INC.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
G06T11/00G06T5/70G06V10/82G06V40/171G06T2207/20081G06T2207/30201G06T2210/32G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,917
App. No.
18/320,857
Filed
May 19, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
2614
USPC
345/629
Abstract

One embodiment of the present invention sets forth a technique for performing identity-preserving image generation. The technique includes converting an identity image depicting a facial identity into an identity embedding. The technique further includes generating a combined embedding based on the identity embedding and a diffusion iteration identifier. The technique further includes converting, using a neural network and based on the combined embedding, a first input image that includes first noise into a first predicted image depicting one or more facial features that include one or more first facial identity features, wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the identity image and are based at least on the identity embedding.

Claims (41)

1 . A computer-implemented method for performing identity-preserving image generation, the computer-implemented method comprising:

converting a first identity image depicting a first facial identity into a first identity embedding;

generating a combined embedding based on the first identity embedding and a diffusion iteration identifier; and

converting, using a neural network and based on the combined embedding, an input image that includes noise into a first predicted image depicting one or more facial features that include one or more first facial identity features,

wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the first identity image and are based at least on the first identity embedding.

2 . The computer-implemented method of claim 1 , further comprising causing the first predicted image to be outputted to a computing device.

3 . The computer-implemented method of claim 1 , wherein at least a portion of the noise is converted to at least a portion of the one or more facial features depicted in the first predicted image.

4 . The computer-implemented method of claim 1 , wherein converting the input image into the first predicted image is repeated for a plurality of diffusion iterations.

5 . The computer-implemented method of claim 4 , wherein the first predicted image generated in each diffusion iteration of the plurality of diffusion iterations includes less noise than a second predicted image generated in a previous diffusion iteration of the plurality of diffusion iterations.

6 . The computer-implemented method of claim 4 , wherein the input image is generated based on a second predicted image that was generated by a previous diffusion iteration of the plurality of diffusion iterations.

7 . The computer-implemented method of claim 4 , wherein the input image is a pure noise image in a first diffusion iteration of the plurality of diffusion iterations.

8 . The computer-implemented method of claim 1 , wherein converting the input image into the first predicted image comprises:

converting, using the neural network and based on the combined embedding, the input image to predicted noise; and

converting the predicted noise to the first predicted image.

9 . The computer-implemented method of claim 8 , further comprising:

training the neural network based on a loss associated with the predicted noise and a pure noise image.

10 . The computer-implemented method of claim 9 , wherein the loss comprises a reconstruction loss determined based on a difference between the predicted noise and the pure noise image.

11 . The computer-implemented method of claim 1 , wherein the one or more identity features include one or more of hair color or gender.

12 . The computer-implemented method of claim 1 , further comprising:

converting one or more attribute values into an attribute embedding,

wherein the combined embedding is further based on the attribute embedding, and

wherein the one or more facial features further include one or more attribute-based facial features, and wherein each attribute-based facial feature is based on one or more of the attribute values.

13 . The computer-implemented method of claim 12 , wherein the one or more attribute-based facial features include one or more of emotion, head pose, glasses, makeup, age, or facial hair.

14 . The computer-implemented method of claim 1 , further comprising:

generating an interpolated identity embedding by interpolating between the first identity embedding and a second identity embedding; and

converting, using the neural network and based on the interpolated identity embedding, the input image that includes the noise into a second predicted image depicting one or more second facial features that include one or more second facial identity features,

wherein the one or more second facial identity features are based on a combination of the first identity image from which the first identity embedding was generated and a second identity image from which the second identity embedding was generated.

15 . The computer-implemented method of claim 1 , further comprising:

identifying, using metadata for a data set on which the diffusion model is trained, a direction in the first identity embedding, wherein the direction corresponds to a facial identity feature;

updating the first identity embedding by adding a determined value to the first identity embedding; and

converting, using a neural network and based on the first identity embedding, the input image that includes the noise into a second predicted image depicting a modified facial identity feature that has changed by an amount that corresponds to the determined value.

16 . The computer-implemented method of claim 1 , further comprising:

identifying a second identity embedding that represents a second facial identity in the latent space;

generating an interpolated identity embedding based on a weighted average of the first identity embedding and the second identity embedding, wherein the first identity embedding is weighted based on a first weight, and the second identity embedding is weighted based on a second weight; and

converting, using a neural network and based on the interpolated identity embedding, the input image that includes the noise into a second predicted image depicting one or more interpolated facial features, wherein the one or more first facial identity features of the first facial identity have the first weight in the interpolated facial features, and one or more second facial identity features of the second facial identity have the second weight in the interpolated facial features.

17 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

converting an identity image depicting a facial identity into an identity embedding;

generating a combined embedding based on the identity embedding and a diffusion iteration identifier; and

converting, using a neural network and based on the combined embedding, a first input image that includes first noise into a first predicted image depicting one or more facial features that include one or more first facial identity features,

wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the identity image and are based at least on the identity embedding.

18 . The one or more non-transitory computer-readable media of claim 17 , further comprising causing the first predicted image to be outputted to a computing device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 063819/0510 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: KANSY, MANUEL JAKOB; RAEL, ANTON JULIEN; NARUNIEC, JACEK KRZYSZTOF; SCHROERS, CHRISTOPHER RICHARD; WEBER, ROMANN MATTHEW
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 063795/0229 →
Continuity (2)
Provisional Application 63343928 · May 19, 2022
Related Publication 20230377214A1 · Nov 23, 2023
References Cited (84)
US 20190238568A1 · Goswami · 2019 [cited by examiner]
US 20200097767A1 · Perry · 2020 [cited by examiner]
US 20220012511A1 · Rowe · 2022 [cited by examiner]
US 20230067841A1 · Saharia · 2023 [cited by examiner]
US 20230222628A1 · Zhao · 2023 [cited by examiner]
Bussi et al., “Accurate Sampling Using Langevin Dynamics”, Physical Review E, vol. 75, DOI: 10.1103/PhysRevE.75.056707, May 25, 2007, pp. 056707-1-056707-7. [cited by applicant]
Chen et al., “SimSwap: An Efficient Framework for High Fidelity Face Swapping”, Deep Learning for Multimedia, In Proceedings of the 28th ACM International Conference on Multimedia, https://doi.org/10.1145/3394171.341363… [cited by applicant]
Choi et al., “StarGAN v2: Diverse Image Synthesis for Multiple Domains”, arXiv:1912.01865, Apr. 26, 2020, 14 pages. [cited by applicant]
Cole et al., “Synthesizing Normalized Faces from Facial Identity Features”, arXiv:1701.04851, Oct. 17, 2017, 14 pages. [cited by applicant]
Creswell et al., “Inverting the Generator of a Generative Adversarial Network”, IEEE Transactions on Neural Networks and Learning Systems, DOI 10.1109/TNNLS.2018.2875194, 2018, pp. 1-8. [cited by applicant]
Deng et al., “RetinaFace: Single-shot Multi-level Face Localisation in the Wild”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, DOI 10.1109/CVPR42600.2020.00525, 2020, pp. 5202-5211. [cited by applicant]
Deng et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, DOI 10.1109/CVPR.2019.00482, 2019, pp. 4685-4694. [cited by applicant]
Dhariwal et al., “Diffusion Models Beat GANs on Image Synthesis”, arXiv:2105.05233, Jun. 1, 2021, pp. 1-44. [cited by applicant]
Duong et al., “Vec2Face: Unveil Human Faces from their Blackbox Features in Face Recognition”, arXiv:2003.06958, Mar. 16, 2020, 10 pages. [cited by applicant]
Efron, Bradley, “Tweedie's Formula and Selection Bias”, Journal of the American Statistical Association, DOI: 10.1198/jasa.2011.tm11181, vol. 106, No. 496, Dec. 2011, pp. 1602-1614. [cited by applicant]
Goodfellow et al., “Explaining and Harnessing Adversarial Examples”, arXiv:1412.6572, Dec. 20, 2014, pp. 1-10. [cited by applicant]
Goswami et al., “Unravelling Robustness of Deep Learning based Face Recognition Against Adversarial Attacks”, Association for the Advancement of Artificial Intelligence, arXiv:1803.00401, Feb. 22, 2018, 8 pages. [cited by applicant]
Guo et al., “Towards Fast, Accurate and Stable 3D Dense Face Alignment”, arXiv:2009.09960, Sep. 21, 2020, pp. 1-21. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Nash Equilibrium”, arXiv:1706.08500, Jun. 26, 2017, pp. 1-61. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models”, 34th Conference on Neural Information Processing Systems, arXiv:2006.11239, Dec. 16, 2020, pp. 1-25. [cited by applicant]
Ho et al., “Classifier-Free Diffusion Guidance”, arXiv:2207.12598, Jul. 26, 2022, pp. 1-14. [cited by applicant]
Huang et al., “Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments”, 2008, pp. 1-11. [cited by applicant]
Karras et al., “Elucidating the Design Space of Diffusion-Based Generative Models”, 36th Conference on Neural Information Processing Systems, arXiv:2206.00364, Oct. 11, 2022, pp. 1-47. [cited by applicant]
Karras et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, arXiv:1812.04948, Mar. 29, 2019, pp. 1-12. [cited by applicant]
Kim et al., “Smooth-Swap: A Simple Enhancement for Face-Swapping with Smoothness”, arXiv:2112.05907, May 5, 2022, pp. 1-16. [cited by applicant]
Kim et al., “Noise2Score: Tweedie's Approach to Self-Supervised Image Denoising without Clean Images”, 35th Conference on Neural Information Processing Systems, arXiv:2106.07009, Oct. 27, 2021, pp. 1-16. [cited by applicant]
Kim et al., “AdaFace: Quality Adaptive Margin for Face Recognition”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1-17. [cited by applicant]
Kortylewski et al., “Empirically Analyzing the Effect of Dataset Biases on Deep Face Recognition Systems”, arXiv:1712.01619, Apr. 19, 2018, pp. 1-11. [cited by applicant]
Kynkäänniemi et al., “Improved Precision and Recall Metric for Assessing Generative Models”, 33rd Conference on Neural Information Processing Systems, arXiv:1904.06991, Oct. 30, 2019, pp. 1-16. [cited by applicant]
Lee et al., “DRIT++: Diverse Image-to-Image Translation via Disentangled Representations”, International Journal of Computer Vision, https://doi.org/10.1007/s11263-019-01284-z, Feb. 3, 2020, 16 pages. [cited by applicant]
Li et al., “Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation”, In Proceedings of the 28th ACM International Conference on Multimedia, https://doi.org/10.1145/3394171.3413960, Oct. 12-16… [cited by applicant]
Liu et al., “Learning to Learn across Diverse Data Biases in Deep Face Recognition”, 2022, pp. 4072-4082. [cited by applicant]
Liu et al., “SphereFace: Deep Hypersphere Embedding for Face Recognition”, arXiv:1704.08063, Apr. 26, 2017, 9 pages. [cited by applicant]
Mahendran et al., “Understanding Deep Image Representations by Inverting Them”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 5188-5196. [cited by applicant]
Mignon et al., “Reconstructing Faces from their Signatures using RBF Regression”, 2013, pp. 1-12. [cited by applicant]
Mohanty et al., “From Scores to Face Templates: A Model-Based Approach”, IEEE Transactions on Pattern Analysis and Machine Intelligence, DOI 10.1109/TPAMI.2007.1129, vol. 29, No. 12, Dec. 2007, pp. 2065-2078. [cited by applicant]
Mordvintsev et al., “Inceptionism: Going Deeper into Neural Networks”, Jun. 18, 2015, pp. 1-7. [cited by applicant]
Moschoglou et al., “AgeDB: The First Manually Collected, In-The-Wild Age Database”, IEEE Conference on Computer Vision and Pattern Recognition Workshops, DOI 10.1109/CVPRW.2017.250, 2017, pp. 1997-2005. [cited by applicant]
Nichol et al., “Improved Denoising Diffusion Probabilistic Models”, arXiv:2102.09672, Feb. 18, 2021, pp. 1-17. [cited by applicant]
Nitzan et al., “Face Identity Disentanglement via Latent Space Mapping”, ACM Trans. Graph, https://doi.org/10.1145/3414685.3417826, vol. 39, No. 6, Article 225, Dec. 2020, pp. 225:1-225:23. [cited by applicant]
Parkhi et al., “Deep Face Recognition”, Visual Geometry Group, 2015, pp. 1-12. [cited by applicant]
Qiu et al., “End2End Occluded Face Recognition by Masking Corrupted Features”, IEEE Transactions on Pattern Analysis and Machine Intelligence, arXiv:2108.09468, Aug. 21, 2021, pp. 1-14. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020, Feb. 26, 2021, pp. 1-48. [cited by applicant]
Ramesh et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents”, arXiv:2204.06125, Apr. 23, 2022, 27 pages. [cited by applicant]
Razzhigaev et al., “Black-Box Face Recovery from Identity Features”, arXiv:2007.13635, Jul. 30, 2020, pp. 1-14. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv:1505.04597, May 18, 2015, 8 pages. [cited by applicant]
Saharia et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, arXiv:2205.11487, May 23, 2022, pp. 1-46. [cited by applicant]
Schroff et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering”, arXiv:1503.03832, Jun. 17, 2015, pp. 1-10. [cited by applicant]
Sengupta et al., “Frontal to Profile Face Verification in the Wild”, 2016, 9 pages. [cited by applicant]
Simonyan et al., “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps”, arXiv:1312.6034, Dec. 20, 2013, pp. 1-8. [cited by applicant]
Singh et al., “On the Robustness of Face Recognition Algorithms Against Attacks and Bias”, Association for the Advancement of Artificial Intelligence, arXiv:2002.02942, Feb. 7, 2020, 7 pages. [cited by applicant]
Song et al., “Score-Based Generative Modeling Through Stochastic Differential Equations”, arXiv:2011.13456, Nov. 26, 2020, pp. 1-32. [cited by applicant]
Sun et al., “Deep Learning Face Representation by Joint Identification-Verification”, arXiv:1406.4773, Jun. 18, 2014, pp. 1-9. [cited by applicant]
Taigman et al., “DeepFace: Closing the Gap to Human-Level Performance in Face Verification”, IEEE Conference on Computer Vision and Pattern Recognition, DOI 10.1109/CVPR.2014.220, 2014, pp. 1701-1708. [cited by applicant]
Terhorst et al., “A Comprehensive Study on Face Recognition Biases Beyond Demographics”, Journal of Latex Class Files, arXiv:2103.01592, vol. 14, No. 8, Mar. 2, 2021, pp. 1-14. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems, arXiv:1706.03762, Dec. 6, 2017, pp. 1-15. [cited by applicant]
Vendrow et al., “Realistic Face Reconstruction from Deep Embeddings”, 35th Conference on Neural Information Processing Systems, 2021, pp. 1-6. [cited by applicant]
Wang et al., “Additive Margin Softmax for Face Verification”, arXiv:1801.05599, May 30, 2018, pp. 1-7. [cited by applicant]
Wang et al., “CosFace: Large Margin Cosine Loss for Deep Face Recognition”, Tencent AI Lab, arXiv:1801.09414, Apr. 3, 2018, 11 pages. [cited by applicant]
Welling et al., “Bayesian Learning via Stochastic Gradient Langevin Dynamics”, In Proceedings of the 28th International Conference on Machine Learning, 2011, 8 pages. [cited by applicant]
Xia et al., “GAN Inversion: A Survey”, arXiv:2101.05278, Mar. 22, 2022, pp. 1-17. [cited by applicant]
Yang et al., “Neural Network Inversion in Adversarial Setting via Background Knowledge Alignment”, Session 2B: ML Security, Association for Computing Machinery, https://doi.org/10.1145/3319535.3354261, Nov. 11-15, 2019,… [cited by applicant]
Yi et al., “Learning Face Representation from Scratch”, arXiv:1411.7923, Nov. 28, 2014, pp. 1-9. [cited by applicant]
Yosinski et al., “Understanding Neural Networks Through Deep Visualization”, 31st International Conference on Machine Learning, arXiv:1506.06579, Jun. 22, 2015, pp. 1-12. [cited by applicant]
Zhang et al., “Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks”, IEEE Signal Processing Letters, DOI 10.1109/LSP.2016.2603342, vol. 23, No. 10, Oct. 2016, pp. 1499-1503. [cited by applicant]
Zhang et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, arXiv:1801.03924, Apr. 10, 2018, pp. 1-14. [cited by applicant]
Zhmoginov et al., “Inverting face embeddings with convolutional neural networks”, arXiv:1606.04189, Jul. 7, 2016, pp. 1-12. [cited by applicant]
Cao et al., “VGGFace2: A Dataset for Recognising Faces Across Pose and Age”, arXiv:1710.08092, Oct. 23, 2017, 10 pages. [cited by applicant]
Che et al., “Mode Regularized Generative Adversarial Networks”, ICLR, arXiv:1612.02136, Dec. 7, 2016, pp. 1-12. [cited by applicant]
Cong et al., “Face Dataset Augmentation with Generative Adversarial Network”, Journal of Physics: Conference Series, doi:10.1088/1742-6596/2218/1/012035, 2022, pp. 1-7. [cited by applicant]
Crispell et al., “Dataset Augmentation for Pose and Lighting Invariant Face Recognition”, arXiv:1704.04326, Apr. 14, 2017, 9 pages. [cited by applicant]
Esler, Tim, “Face Recognition Using Pytorch”, Facenet Pytorch, Retrieved from https://github.com/timesler/facenet-pytorch, 2017, pp. 1-9. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets”, arXiv:1406.2661, Jun. 10, 2014, pp. 1-9. [cited by applicant]
He et al., “Arbitrary Facial Attribute Editing: Only Change What You Want”, arXiv:1711.10678, Nov. 29, 2017, pp. 1-9. [cited by applicant]
Li et al., “Generate Identity-Preserving Faces by Generative Adversarial Networks”, arXiv:1706.03227, Jun. 25, 2017, pp. 1-9. [cited by applicant]
Mai et al., “Face Image Reconstruction from Deep Templates”, arXiv:1703.00832, Mar. 2, 2017, pp. 1-14. [cited by applicant]
Masi et al., “Do We Really Need to Collect Millions of Faces for Effective Face Recognition?”, arXiv:1603.07057, Apr. 11, 2016, pp. 1-18. [cited by applicant]
Serengil et al., “LightFace: A Hybrid Deep Face Recognition Framework”, IEEE, 2020, 5 pages. [cited by applicant]
Shorten et al., “A survey on Image Data Augmentation for Deep Learning”, Journal of Big Data, https://doi.org/10.1186/s40537-019-0197-0, vol. 6, 2019, pp. 1-48. [cited by applicant]
Wang et al., “A Survey on Face Data Augmentation”, arXiv:1904.11685, Apr. 26, 2019, 26 pages. [cited by applicant]
Yang et al., “Image Data Augmentation for Deep Learning: A Survey”, arXiv:2204.08610, Apr. 19, 2022, 8 pages. [cited by applicant]
Zeno et al., “PFA-GAN: Pose Face Augmentation Based on Generative Adversarial Network”, Informatica, DOI: https://doi.org/10.15388/21-INFOR443, Jan. 2021, pp. 1-16. [cited by applicant]
Serengil et al., “HyperExtended LightFace: A Facial Attribute Analysis Framework”, 2021 International Conference on Engineering and Emerging Technologies (ICEET). IEEE, https://doi.org/10.1109/ICEET53442.2021.9659697, 4… [cited by applicant]