IP Library Granted Patent US 12,450,705
Granted Patent B2
US 12,450,705 · App. 17/127,506 · Granted Oct 21, 2025

Altering a facial identity in a video stream

Inventors: Oran Gafni (Ramat Gan, IL); Lior Wolf (Herzliya, IL)
Assignee: Meta Platforms, Inc.
G06T5/75G06N20/00G06T5/70G06V40/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,705
App. No.
17/127,506
Granted
Oct 21, 2025
Kind
B2
Abstract

A computing device performs generating a first identity encoding representing a first facial identity of the person based on an image of a person, generating a second identity encoding representing a second facial identity different from the first facial identity of the person based on the first identity encoding, generating a source encoding by using an encoder to process a source image of the person having an expression, generating an intermediate image by using a decoder to process the source encoding and the second identity encoding, the intermediate image including a face having the second facial identity and the expression of the person in the source image, and generating an output image by blending the source image with facial features of the face in the intermediate image.

Claims (59)

1. A method comprising:

generating, based on an image of a person, a first identity encoding representing a first facial identity of the person;

generating a second identity encoding representing a second facial identity different from the first facial identity of the person by processing the first identity encoding with a first facial-identity altering machine-learning model comprising one or more fully connected layers and training data to facilitate generating the second identity encoding, wherein the training data comprises (i) pairs of original face images and corresponding altered face images, and (ii) pairs of images of a different first person and a second person;

generating a source encoding by using an encoder to process a source image of the person comprising an expression;

generating an intermediate image by using a decoder to process the source encoding and the second identity encoding, the intermediate image comprising a face comprising the second facial identity and the expression of the person in the source image; and

generating an output image by blending the source image with facial features of the face in the intermediate image.

2. The method of claim 1 , wherein a first machine-learning model is associated with the encoder and the decoder.

3. The method of claim 2 , further comprising:

utilizing the pairs of images as training data at an iteration of a training procedure associated with the first machine-learning model, wherein the pairs of images comprise a first image of the first person comprising a second expression and a second image of the second person.

4. The method of claim 3 , wherein the iteration of the training procedure comprises:

generating a training source encoding by using the encoder of the first machine-learning model to process a distorted first image of the pair;

generating a training identity encoding by using a pre-trained face recognition machine-learning model to process the second image of the pair;

generating a training intermediate image by using the decoder of the first machine-learning model to process the training source encoding and the training identity encoding;

generating a training output image by blending the first image with facial features of a face in the training intermediate image;

determining losses based on the training intermediate image or the training output image, wherein the losses comprise perceptual losses; and

updating trainable variables of the first machine-learning model based on the computed losses.

5. The method of claim 4 , wherein the distorted first image prevents the first machine-learning model from generating the training intermediate image identical to the first image.

6. The method of claim 4 , further comprising:

determining the perceptual losses at k different layers, and wherein the perceptual losses at i lowest layers out of the k layers are determined based on comparisons between the first image and the training output image, and wherein the perceptual losses at the remaining k−i layers are determined based on comparisons between the second image and the training output image.

7. The method of claim 6 , further comprising:

updating the trainable variables of the first machine-learning model to minimize the perceptual losses.

8. The method of claim 1 , further comprising:

generating the first identity encoding by processing the image with a pre-trained face recognition machine-learning model.

9. The method of claim 1 , further comprising:

training the first facial-identity altering machine-learning model with the pairs of original face images and the corresponding altered face images as the training data.

10. The method of claim 9 , further comprising:

generating an altered face image corresponding to an original face image by processing the original face image with the first facial-identity altering machine-learning model.

11. The method of claim 10 , wherein the first facial-identity altering machine-learning model comprises a second encoder and a second decoder.

12. The method of claim 1 , wherein the source image corresponds to a frame of a video stream.

13. The method of claim 1 , further comprising:

generating a blending mask in an instance in which the source encoding and the second identity encoding are processed by the decoder, and wherein the blending mask represents a blending weight to be applied to the intermediate image at one or more pixels of the output image.

14. The method of claim 13 , wherein blending the source image with facial features of the face in the intermediate image comprises:

creating one or more output images by applying an inverse of the blending mask to the source image; and

projecting the blending mask applied to the intermediate image to the one or more output images.

15. A non-transitory computer-readable medium storing instructions that, when executed, cause:

generating, based on an image of a person, a first identity encoding representing a first facial identity of the person;

generating a second identity encoding representing a second facial identity different from the first facial identity of the person by processing the first identity encoding with a first facial-identity altering machine-learning model comprising one or more fully connected layers and comprising training data to facilitate generating the second identity encoding, wherein the training data comprises (i) pairs of original face images and corresponding altered face images, and (ii) pairs of images of a different first person and a second person;

generating a source encoding by using an encoder to process a source image of the person comprising an expression;

generating an intermediate image by using a decoder to process the source encoding and the second identity encoding, the intermediate image comprising a face comprising the second facial identity and the expression of the person in the source image; and

generating an output image by blending the source image with facial features of the face in the intermediate image.

16. The computer-readable medium of claim 15 , wherein a first machine-learning model is associated with the encoder and the decoder.

17. The computer-readable medium of claim 16 , wherein the instructions, when executed, further cause:

utilizing the pairs of images as training data at an iteration of a training procedure associated with the first machine-learning model, wherein the pairs of images comprise a first image of the first person comprising a second expression and a second image of the second person.

18. The computer-readable medium of claim 17 , wherein the iteration of the training procedure comprises:

generating a training source encoding by using the encoder of the first machine-learning model to process a distorted first image of the pair;

generating a training identity encoding by using a pre-trained face recognition machine-learning model to process the second image of the pair;

generating a training intermediate image by using the decoder of the first machine-learning model to process the training source encoding and the training identity encoding;

generating a training output image by blending the first image with facial features of a face in the training intermediate image;

determining losses based on the training intermediate image or the training output image, wherein the losses comprise perceptual losses; and

updating trainable variables of the first machine-learning model based on the computed losses.

19. A system comprising:

one or more processors; and

at least one non-transitory memory comprising instructions executable by the one or more processors, the one or more processors operable when executing the instructions to:

generate, based on an image of a person, a first identity encoding representing a first facial identity of the person;

generate a second identity encoding representing a second facial identity different from the first facial identity of the person by processing the first identity encoding with a first facial-identity altering machine-learning model comprising one or more fully connected layers and comprising training data to facilitate generating the second identity encoding, wherein the training data comprises (i) pairs of original face images and corresponding altered face images, and (ii) pairs of images of a different first person and a second person;

generate a source encoding by using an encoder to process a source image of the person comprising an expression;

generate an intermediate image by using a decoder to process the source encoding and the second identity encoding, the intermediate image comprising a face comprising the second facial identity and the expression of the person in the source image; and

generate an output image by blending the source image with facial features of the face in the intermediate image.

20. The system of claim 19 , wherein the encoder and the decoder belong to a first machine-learning model.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2020
From: GAFNI, ORAN; WOLF, LIOR
To: FACEBOOK, INC.
Reel/Frame 054763/0773 →
Continuity (1)
Related Publication 20220198617A1 · Jun 23, 2022
References Cited (68)
US 20030001838A1 · Han et al. · 2003 [cited by applicant]
US 20130208005A1 · Kasahara et al. · 2013 [cited by applicant]
US 20180365874A1 · Hadap et al. · 2018 [cited by applicant]
US 20200234482A1 · Krokhalev et al. · 2020 [cited by applicant]
CN 110969572A · 2020 [cited by applicant]
CN 111971713A · 2020 [cited by applicant]
CN 112036219A · 2020 [cited by applicant]
“Oran Gafni et. al., Live Face De-Identification in Video, International Conference on Computer Vision ICCV, Nov. 2019, pp. 9378-9387” (Year: 2019). [cited by examiner]
“Yuval Nirkin et. al., On Face Segmentation, Face Swapping, and Face Perception, 2018 13th IEEE International Conference on Automatic Face and Gesture Recognition” (Year: 2018). [cited by examiner]
“Naser Damer et. al., Deep Learning-based Face Recognition and the Robustness to Perspective Distortion,” Aug. 2018, 2018 International Conference on Pattern Recognition, Beijing, China (Year: 2018). [cited by examiner]
“Aparna Bharati et. al., Detecting Facial Retouching Using Supervised Deep Learning, Nov. 2016, IEEE Transactions on Information Forensics and Security vol. 11, No. 9” (Year: 2016). [cited by examiner]
“Aparna Bharati et. al., Detecting Facial Retouching Using Supervised Deep Learning, Sep. 2016, IEEE Transactions on Information Forensics and Security, vol. 11, Issue 9” (Year: 2016). [cited by examiner]
“Apama Bharati et. al., Detecting Facial Retouching Using Supervised Deep Learning, Sep. 2016, IEEE Transactions on Information Forensics and Security, vol. 11, Issue 9” (Year: 2016). [cited by examiner]
Benaim, et al., One-Sided Unsupervised Domain Mapping, arXiv:1706.00826v2 [cs.CV], In NIPS, 18 pages, Nov. 18, 2017. [cited by applicant]
Bitouk, et al., Face Swapping: Automatically Replacing Faces in Photographs, ACM Transactions on Graphics, vol. 27, No. 3, Article 39, 8 pages, Aug. 2008. [cited by applicant]
Blanz, et al., Exchanging Faces in Images, EUROGRAPHICS 2004 / M.-P. Cani and M. Slater, 23(3):1-8, 2004. [cited by applicant]
Cao, et al., VGGFace2: A Dataset for Recognising Faces Across Pose and Age, arXiv:1710.08097v2 [cs.CV], IFEE, 11 pages, May 13, 2018. [cited by applicant]
Chen, et al., InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets, arXiv:1606.03657v1, 14 pages, Jun. 12, 2016. [cited by applicant]
Chollet, Xception: Deep Learning with Depthwise Separable Convolutions, CVF, pp. 1251-1258. [cited by applicant]
Deng, et al., ArcFace: Additive Angular Margin Loss for Deep Face Recognition, CVPR, Computer Vision Foundation, pp. 4690-4699. [cited by applicant]
Deepfaks/Faceswap, GitHub—Deepfakes/Faceswap: Deepfakes Software for All, available via https://github.com/deepfakes/faceswap, 11 pages, Apr. 6, 2021. [cited by applicant]
Gafni, et al., Live Face De-Identification in Video, ICCV, CVF, pp. 9378-9387. [cited by applicant]
Goodfellow, et al., Generative Adversarial Nets Generative Adversarial Nets, arXiv:1406.2661v1, 9 pages, Jun. 10, 2014. [cited by applicant]
Gross, et al., Semi-Supervised Learning of Multi-Factor Models for Face De-Identification, 8 pages. [cited by applicant]
He, et al., Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification, ICCV, CVF, pp. 1026-1034, Feb. 2015. [cited by applicant]
He, et al., Deep Residual Learning for Image Recognition, CVPR, CVF, pp. 770-778. [cited by applicant]
Huang, et al., Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments, Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition, Erik Learned-Miller and A… [cited by applicant]
Huang, et al., Multimodal Unsupervised Image-to-Image Translation, ECCV, 18 pages, 2018. [cited by applicant]
Johnson, et al., Perceptual losses for real-time style transfer and super-resolution, In ECCV, 18 pages, Mar. 27, 2016. [cited by applicant]
Jourabloo, et al., Attribute Preserved Face Deidentification, Department of Computer Science and Engineering, Michigan State University, 8 pages. [cited by applicant]
Karras, et al., Progressive Growing of Gans for Improved Quality, Stability, and Variation, arXiv:1710.10196v3, In ICLR, 26 pages, Feb. 26, 2018. [cited by applicant]
Kazemi, et al., One Millisecond Face Alignment with an Ensemble of Regression Trees, In Proceedings of the IEEE conference on computer vision and pattern recognition, 8 pages, 2014. [cited by applicant]
Kemelmacher-Shlizerman, Transfiguring Portraits, ACM Trans. Graph., 35(4), Jul. 2016. [cited by applicant]
Kim, et al., Learning to Discover Cross-Domain Relations with Generative Adversarial Networks, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 9 pages, 2017. [cited by applicant]
King, et al., Dlib-ml: A Machine Learning Toolkit, Journal of Machine Learning Research, 10(2009):1755-1758, 2009. [cited by applicant]
Kingma, et al., Adam: A Method for Stochastic Optimization, In ICLR, 15 pages, 2015. [cited by applicant]
Korshunova, et al., Fast Face-Swap Using Convolutional Neural Networks, In the IEEE International Conference on Computer Vision, pp. 3677-3685. [cited by applicant]
Kumar, et al., Attribute and Simile Classifiers for Face Verification, In CVPR, pp. 365-372, 2009. [cited by applicant]
Lample, et al. Fader Networks: Manipulating Images by Sliding Attributes, In NIPS, 10 pages, 2017. [cited by applicant]
Lee, et al., Diverse Image-To-Image Translation Via Disentangled Representations, In the European Conference on Computer Vision (ECCV), 17 pages, 2018. [cited by applicant]
Liu, et al., Unsupervised Image-To-Image Translation Networks, In NIPS, 11 pages, 2017. [cited by applicant]
Liu, et al., Deep Learning Face Attributes in The Wild, In ICCV, pp. 3730-3738. [cited by applicant]
Makhzani, et al., Adversarial Autoencoders, arXiv preprint arXiv: 1511.05644v2, 16 pages, May 25, 2016. [cited by applicant]
Mao, et al., Least Squares Generative Adversarial Networks, In ICCV, pp. 2794-2802. [cited by applicant]
Meden, et al., Face Deidentification with Generative Deep Neural Networks, IET Signal Processing, 11(9):1046-1054, 2017. [cited by applicant]
Newton, et al., Preserving Privacy by De-Identifying Face Images, IEEE transactions on Knowledge and Data Engineering, Carnegie Mellon University, 26 pages, Mar. 2003. [cited by applicant]
Nirkin, et al., On Face Seg-Mentation, Face Swapping, And Face Perception, Arxiv:1704.06729v1, 14 pages, Apr. 22, 2017. [cited by applicant]
Oh, et al., Adversarial Image Perturbation for Privacy Protection a Game Theory Perspective, in 2017 IEEE international conference on computer vision (ICCV), 17 pages, Jul. 26, 2017. [cited by applicant]
Phillips, et al., An Introduction to the Good, the Bad, & the Ugly Face Recognition Challenge Problem, In Automatic Face & Gesture Recognition, 10 pages, Mar. 2011. [cited by applicant]
Radford, et al., Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks, ICLR 2016, 16 pages, 2016. [cited by applicant]
Ronneberger, et al., U-net: Convolutional Networks for Biomedical Image Segmentation, arXiv:1505.04597v1, 8 pages, May 18, 2015. [cited by applicant]
Rössler, et al., FaceForensics: A large-scale Video Dataset for Forgery Detection in Human Faces, arXiv:1803.09179v1, 21 pages, Mar. 24, 2017. [cited by applicant]
Salimans, et al., Improved Techniques for Training Gans, arXiv:1606.03498v1, 10 pages, Jun. 10, 2016. [cited by applicant]
Samarzija, et al., An Approach to the De-Identification of Faces in Different Poses, University of Zagreb, 13 pages, May 29, 2014. [cited by applicant]
Schroff, et al., Facenet: A Unified Embedding for Face Recognition and Clustering, CVPR, pp. 815-823, 2015. [cited by applicant]
Sun, et al., Natural and Effective Obfuscation by Head Inpainting, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5050-5059. [cited by applicant]
Sun, et al., A hybrid Model for Identity Obfuscation by Face Replacement, In Proceedings of the European Conference on Computer Vision (ECCV), 17 pages, 2018. [cited by applicant]
Taigman, et al., Unsupervised Cross-Domain Image Generation, In International Conference on Learning Representations (ICLR), 14 pages, 2017. [cited by applicant]
Thies, et al., Face2face: Real-time Face Capture and Reenactment of RGB Videos, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2387-2395. [cited by applicant]
Ulyanov, et al., Texture networks: Feed-forward Synthesis of Textures and Stylized Images, In ICML, 9 pages, 2016. [cited by applicant]
Ulyanov, et al., Instance Normalization: The Missing Ingredient for Fast Stylization, arXiv preprint arXiv:1607.08022v3, 6 pages, Nov. 6, 2017. [cited by applicant]
Wu, et al., Privacy-Protective-Gan For Face De-Identification, arXiv preprint arXiv:1806.08906v1, 11 pages, Nov. 23, 2018. [cited by applicant]
Xie, et al., Feature Denoising for Improving Adversarial Robustness, CVPR paper, 9 pages. [cited by applicant]
Yi, et al., DualGAN: Unsupervised Dual Learning for Image-To-Image Translation, ICCV pager, 9 pages. [cited by applicant]
Zhang, et al., Mixup: Beyond Empirical Risk Minimization, ICLR, arXiv:1710.09412v2, 13 pages, Apr. 27, 2018. [cited by applicant]
Huang P., et al., “Learning Identity-Invariant Motion Representations for Cross-ID Face Reenactment,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13, 2020, pp. 7084-7092. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2021/063372 mailed Apr. 4, 2022, 9 pages. [cited by applicant]
Office Action mailed Jul. 1, 2025 for Chinese Application No. 202180094210.X, filed Dec. 14, 2021, 8 pages. [cited by applicant]