IP Library › Granted Patent US 12,688,325
Granted Patent B2
US 12,688,325 · App. 18/224,916 · Granted Jul 21, 2026

Face anonymization in digital images

Inventors: Yang Yang (Santa Clara, CA); Zhixin Shu (San Jose, CA); Shabnam Ghadar (Menlo Park, CA); Jingwan Lu (Santa Clara, CA); Jakub Fiser (Seattle, WA); Elya Schechtman (Seattle, WA); Cameron Y. Smith (Santa Cruz, CA); Baldo Antonio Faieta (San Francisco, CA); Alex Charles Filipkowski (San Francisco, CA)
Assignee: Adobe Inc.
G06T11/60G06F16/532G06F16/56G06F21/6254G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,325
App. No.
18/224,916
Filed
Jul 21, 2023
Granted
Jul 21, 2026
Kind
B2
Art Unit
2673
USPC
382/233
Abstract

Face anonymization techniques are described that overcome conventional challenges to generate an anonymized face. In one example, a digital object editing system is configured to generate an anonymized face based on a target face and a reference face. As part of this, the digital object editing system employs an encoder as part of machine learning to extract a target encoding of the target face image and a reference encoding of the reference face. The digital object editing system then generates a mixed encoding from the target and reference encodings. The mixed encoding is employed by a machine-learning model of the digital object editing system to generate a mixed face. An object replacement module is used by the digital object editing system to replace the target face in the target digital image with the mixed face.

Claims (44)

1 . A method comprising:

receiving, by a processing device, a target digital image depicting a target face;

automatically editing, by the processing device, the target digital image to remove a portion of the target digital image around the target face;

encoding, by the processing device using a machine learning model, data related to the target digital image to generate a target encoding including feature vectors representing three-dimensional aspects of a pose of the target face;

generating, by the processing device using the machine learning model, data including a descriptor of the pose of the target face based on the target encoding;

generating, by the processing device, a search query for a reference face in a reference digital image based on the descriptor of the pose of the target face;

automatically selecting, by the processing device, the reference digital image from candidate digital image results resulting from the search query by comparing the feature vectors representing the three-dimensional aspects of the pose of the target face and feature vectors representing three-dimensional aspects of poses of candidate faces of the candidate digital image results; and

generating, by the processing device, a mixed face by editing the target digital image to incorporate a portion of the reference face.

2 . The method of claim 1 , wherein the descriptor of the pose of the target face indicates a facial pose or a facial feature.

3 . The method of claim 1 , further comprising cropping the reference face from the reference digital image.

4 . The method of claim 1 , further comprising extracting the target encoding from the target face and a reference encoding from the reference face.

5 . The method of claim 4 , wherein the mixed face is generated by mixing the target encoding with the reference encoding.

6 . The method of claim 5 , wherein the target encoding is mixed with the reference encoding using linear interpolation.

7 . The method of claim 1 , further comprising replacing the target face in the target digital image with the mixed face.

8 . The method of claim 1 , wherein automatically selecting the reference digital image further comprises comparing a three-dimensional target face shape model with a three-dimensional reference face shape model.

9 . The method of claim 1 , further comprising encoding, using the machine learning model, the candidate digital image results to generate the feature vectors representing the three-dimensional aspects of the poses of the candidate faces of the candidate digital image results.

10 . The method of claim 1 , wherein a similarity score defines a level of visual similarity between the pose of the target face and the poses of the candidate faces of the candidate digital image results for automatically selecting the reference digital image.

11 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving a target digital image depicting a target object;

automatically editing the target digital image to remove a portion of the target digital image around the target object;

encoding, using a machine learning model, data related to the target digital image to generate a target encoding including feature vectors representing three-dimensional aspects of a pose of the target object;

generating, using the machine learning model, data including a descriptor of the pose of the target object based on the target encoding;

generating a search query for a reference object in a reference digital image based on the descriptor of the pose of the target object;

executing a search of digital images using the search query;

automatically selecting the reference digital image from candidate digital image results resulting from the search query by comparing the feature vectors representing the three-dimensional aspects of the pose of the target object and feature vectors representing three-dimensional aspects of poses of candidate objects of the candidate digital image results; and

generating a mixed object by editing the target digital image to incorporate a portion of the reference object.

12 . The system of claim 11 , wherein the descriptor of the pose of the target object indicates a feature of the target object.

13 . The system of claim 11 , further comprising extracting the target encoding from the target object and a reference encoding from the reference object.

14 . The system of claim 13 , wherein the mixed object is generated by mixing the target encoding with the reference encoding.

15 . The system of claim 14 , wherein the target encoding is mixed with the reference encoding using linear interpolation.

16 . The system of claim 11 , further configured to perform operations comprising encoding, using the machine learning model, the candidate digital image results to generate the feature vectors representing the three-dimensional aspects of the poses of the candidate objects of the candidate digital image results.

17 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving a target digital image depicting a target face;

automatically editing the target digital image to remove a portion of the target digital image around the target face;

encoding, using a machine learning model, data related to the target digital image to generate a target encoding including feature vectors representing three-dimensional aspects of a pose of the target face;

generating, using the machine learning model, data including a descriptor of the pose of the target face based on the target encoding;

generating a search query for a reference face in a reference digital image based on the descriptor of the pose of the target face;

automatically selecting the reference digital image from candidate digital image results resulting from the search query by comparing the feature vectors representing the three-dimensional aspects of the pose of the target face and feature vectors representing three-dimensional aspects of poses of candidate faces of the candidate digital image results; and

generating a mixed face by editing the target digital image to incorporate a portion of the reference face.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the descriptor of the pose of the target face indicates a facial pose or a facial feature.

19 . The non-transitory computer-readable storage medium of claim 17 , further comprising cropping the reference face from the reference digital image.

20 . The non-transitory computer-readable storage medium of claim 17 , further configured to perform operations comprising encoding, using the machine learning model, the candidate digital image results to generate the feature vectors representing the three-dimensional aspects of the poses of the candidate faces of the candidate digital image results.

Continuity (2)
Continuation 17094093 · Nov 10, 2020
Related Publication 20230360299A1 · Nov 9, 2023
References Cited (29)
US 11106919B1 · Balogh · 2021 [cited by examiner]
US 11748928B2 · Yang et al. · 2023 [cited by applicant]
US 20180314878A1 · Lee · 2018 [cited by examiner]
US 20190012442A1 · Hunegnaw · 2019 [cited by examiner]
US 20190122329A1 · Wang · 2019 [cited by examiner]
US 20200033615A1 · Kim · 2020 [cited by examiner]
US 20200134858A1 · Yang · 2020 [cited by examiner]
US 20200186721A1 · Ogawa · 2020 [cited by examiner]
US 20220012362A1 · Kuta · 2022 [cited by examiner]
US 20220148243A1 · Yang et al. · 2022 [cited by applicant]
US 20230138380A1 · Chen · 2023 [cited by examiner]
Muench et al, Data Anonymization for Data Protection on Publicly Recorded Data, Nov. 2019, Lecture Notes in Computer Science (Year: 2019). [cited by examiner]
T. Kim and J. Yang, “Latent-Space-Level Image Anonymization With Adversarial Protector Networks,” in IEEE Access, vol. 7, pp. 84992-84999, 2019, (Year: 2019). [cited by examiner]
Berretti, Stefano, Alberto Del Bimbo, and Pietro Pala. “3D partial face matching using local shape descriptors.” Proceedings of the 2011 joint ACM workshop on Human gesture and behavior understanding. 2011. (Year: 2011). [cited by examiner]
“DeepFakes_Faceswap”, GitHub.com, deepfakes [retrieved Dec. 10, 2020]. Retrieved from the Internet <https://github.com/deepfakes/faceswap>., Feb. 25, 2018, 11 pages. [cited by applicant]
“Eigen V3”, TuxFamily.org [retrieved Feb. 10, 2021]. Retrieved from the Internet <https://eigen.tuxfamily.org/index.php?title=Main_Page>., Nov. 2019, 13 pages. [cited by applicant]
U.S. Appl. No. 17/094,093 , “First Action Interview Office Action”, U.S. Appl. No. 17/094,093, filed Apr. 5, 2023, 4 pages. [cited by applicant]
U.S. Appl. No. 17/094,093 , “Notice of Allowance”, U.S. Appl. No. 17/094,093, filed May 22, 2023, 7 pages. [cited by applicant]
U.S. Appl. No. 17/094,093 , “Pre-Interview First Office Action”, U.S. Appl. No. 17/094,093, filed Mar. 1, 2023, 3 pages. [cited by applicant]
Bitouk, Dmitri , et al., “Face Swapping: Automatically Replacing Faces in Photographs”, ACM SIGGRAPH 2008 papers [retrieved Dec. 15, 2020]. Retrieved from the Internet <http://citeseerx.ist.psu.edu/viewdoc/download?doi=… [cited by applicant]
Bradski, Gary , et al., “The OpenCV Library”, Dr. Dobb's, The World of Software Development Blog [retrieved Mar. 16, 2023]. Retrieved from the Internet <https://www.drdobbs.com/open-source/the-opencv-library/184404319>.… [cited by applicant]
Deng, Jiankang , et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition”, Cornell University arXiv, arXiv.org [retrieved Feb. 20, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1801.07698.pd… [cited by applicant]
He, Kaiming , et al., “Deep Residual Learning for Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition, 2016 [retrieved Feb. 18, 2022], Retrieved from the Internet: <https://… [cited by applicant]
Karras, Tero , et al., “A Style-Based Generator Architecture for Generative Adversarial Networks”, Cornell University arXiv, arXiv.org [retrieved Jun. 28, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1812.0… [cited by applicant]
Kemelmacher-Shlizerman, Ira , “Transfiguring Portraits”, ACM Transactions on Graphics (TOG) 35, No. 4 [retrieved Dec. 15, 2020]. Retrieved from the Internet <https://homes.cs.washington.edu/~kemelmi/Transfiguring_Portra… [cited by applicant]
Naruniec, Jacek , et al., “High-Resolution Neural Face Swapping for Visual Effects”, Eurographics Symposium on Rendering 2020, vol. 39, No. 4 [retrieved Dec. 10, 2020]. Retrieved from the Internet <https://studios.disne… [cited by applicant]
Sun, Ke , et al., “Deep High-Resolution Representation Learning for Human Pose Estimation”, arXiv Preprint, arXiv.org [retrieved Feb. 10, 2021]. Retrieved from the Internet <https://arxiv.org/pdf/1902.09212.pdf>., Feb. … [cited by applicant]
Thies, Justus , et al., “Face2Face: Real-Time Face Capture and Reenactment of RGB Videos”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016 [retrieved Feb. 10, 2021]. Retrieved … [cited by applicant]
Yang, Fei , et al., “Expression Flow for 3D-Aware Face Component Transfer”, ACM Trans. Graph. 30, 4, Article 60 [retrieved Feb. 10, 2021]. Retrieved from the Internet <https://citeseerx.ist.psu.edu/viewdoc/download?doi=… [cited by applicant]