IP Library › Granted Patent US 12,462,560
Granted Patent B2
US 12,462,560 · App. 17/845,353 · Granted Nov 4, 2025

Video manipulation detection

Inventors: Ritwik Sinha (Cupertino, CA); Viswanathan Swaminathan (Saratoga, CA); Trisha Mittal (College Park, MD); John Philip Collomosse (Pyrford, GB)
Assignee: Adobe Inc.
G06V20/41G06T7/0002G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,560
App. No.
17/845,353
Granted
Nov 4, 2025
Kind
B2
Abstract

Techniques for video manipulation detection are described to detect one or more manipulations present in digital content such as a digital video. A detection system, for instance, receives a frame of a digital video that depicts at least one entity. Coordinates of the frame that correspond to a gaze location of the entity are determined, and the detection system determines whether the coordinates correspond to a portion of an object depicted in the frame to calculate a gaze confidence score. A manipulation score is generated that indicates whether the digital video has been manipulated based on the gaze confidence score. In some examples, the manipulation score is based on at least one additional confidence score.

Claims (41)

1 . A method comprising:

receiving, in a user interface of a processing device, a frame of a digital video that depicts an entity within a scene;

identifying, by the processing device, coordinates of the frame that correspond to a gaze location within the scene of the entity included in the scene using a gaze tracking algorithm;

determining, by the processing device, whether the coordinates correspond to a location of an object depicted in the scene;

calculating, by the processing device, a gaze confidence score that indicates a probability that the scene includes an added or removed visual feature based on the determining;

generating, by the processing device, a manipulation score that indicates whether the digital video has been manipulated based on the gaze confidence score and at least one additional confidence score; and

outputting, by the processing device, the manipulation score in the user interface.

2 . The method as described in claim 1 , further comprising:

determining that the digital video has been manipulated based on the manipulation score; and

generating an indication of a spatial location of a manipulation in the digital video for display in the user interface.

3 . The method as described in claim 1 , wherein the at least one additional confidence score includes a visual artifact confidence score calculated by leveraging a convolutional neural network to detect resolution inconsistencies associated with affine warpings.

4 . The method as described in claim 1 , wherein the at least one additional confidence score includes a temporal confidence score calculated by determining an optical flow between adjacent frames of a plurality of frames of the digital video.

5 . The method as described in claim 4 , wherein the temporal confidence score is calculated by comparing audial features of the digital video to visual features of the digital video.

6 . The method as described in claim 1 , wherein the at least one additional confidence score includes an affective state confidence score calculated by leveraging a machine learning model to detect discrepancies in affective states for the entity.

7 . The method as described in claim 6 , wherein the affective states are partially based on facial expressions of the entity, body postures of the entity, or contextual data associated with the digital video.

8 . The method as described in claim 1 , wherein the manipulation score is a maximum of the gaze confidence score and the at least one additional confidence score.

9 . The method as described in claim 1 , wherein generating the manipulation score includes applying a weighting to at least one of the gaze confidence score, a visual artifact confidence score, a temporal confidence score, or an affective state confidence score.

10 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device configured to perform operations including:

receiving, in a user interface of the processing device, a frame of a digital video that depicts an individual within a scene;

identifying coordinates of the frame that correspond to a gaze location within the scene of the individual using a gaze tracking algorithm;

determining whether the coordinates correspond to a portion of an object depicted in the scene;

calculating a gaze confidence score based on the determining that indicates a probability that the scene includes one or more manipulated digital objects; and

generating a manipulation score for output in the user interface that indicates whether the digital video has been manipulated based on the gaze confidence score and one or more of a visual artifact confidence score, a temporal confidence score, or an affective state confidence score.

11 . The system as described in claim 10 , the operations further including:

determining that the digital video is manipulated based on the manipulation score being greater than a threshold; and

generating an indication of a type of manipulation present in the digital video for display in the user interface.

12 . The system as described in claim 10 , wherein the visual artifact confidence score is calculated by leveraging a convolutional neural network to detect resolution inconsistencies associated with affine warpings.

13 . The system as described in claim 10 , wherein the temporal confidence score is calculated by determining an optical flow between adjacent frames of a plurality of frames of the digital video.

14 . The system as described in claim 10 , wherein the temporal confidence score is calculated by comparing audial features of the digital video to visual features of the digital video.

15 . The system as described in claim 10 , wherein the digital video includes a plurality of individuals and wherein the affective state confidence score is calculated by leveraging a machine learning model to generate labels for each individual of the plurality of individuals, and comparing the labels, one to another.

16 . The system as described in claim 10 , wherein the manipulation score is a maximum of the gaze confidence score and the one or more of the visual artifact confidence score, the temporal confidence score, or the affective state confidence score.

17 . A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving, in a user interface of the processing device, a frame of a digital video that depicts an entity within a scene;

identifying gaze location coordinates of the scene that correspond to a gaze location of the entity included in the scene using a gaze tracking algorithm;

calculating a gaze confidence score based on whether the gaze location coordinates correspond to a portion of an object depicted in the scene, the gaze confidence score indicating a probability that the digital video includes one or more manipulated digital objects within the scene; and

presenting a manipulation score that indicates whether the digital video has been manipulated based on the gaze confidence score and at least one additional confidence score.

18 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the manipulation score is based on an aggregation of the gaze confidence score and at least one of a visual artifact confidence score, a temporal confidence score, or an affective state confidence score.

19 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the frame depicts a plurality of entities within the scene, and wherein calculating the gaze confidence score includes comparing a plurality of gaze locations within the scene for each entity of the plurality of entities, one to another.

20 . The non-transitory computer-readable storage medium as described in claim 17 , wherein the one or more manipulated digital objects include one or more of an object added to the scene or an object that has been removed from the scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2022
From: SINHA, RITWIK; SWAMINATHAN, VISWANATHAN; MITTAL, TRISHA; COLLOMOSSE, JOHN PHILIP
To: ADOBE INC.
Reel/Frame 060263/0798 →
Continuity (1)
Related Publication 20230410505A1 · Dec 21, 2023
References Cited (32)
US 10546193B2 · Schmidt · 2020 [cited by examiner]
US 20200134295A1 · el Kaliouby · 2020 [cited by examiner]
US 20220138472A1 · Mittal · 2022 [cited by examiner]
US 20220269922A1 · Mathews · 2022 [cited by examiner]
Li, Y. “Exposing deepfake videos by detecting face warping artifacts.” arXiv preprint arXiv:1811.00656 (2018) (Year: 2018). [cited by examiner]
Mehta, Vineet, et al. “Fakebuster: a deepfakes detection tool for video conferencing scenarios.” Companion Proceedings of the 26th International Conference on Intelligent User Interfaces. 2021. (Year: 2021). [cited by examiner]
Prajapati, Pratikkumar. “Quantifying DeepFake Detection Accuracy for a Variety of Natural Settings.” (2020) (Year: 2020). [cited by examiner]
Li, Xiaodan, et al. “Sharp multiple instance learning for deepfake video detection.” Proceedings of the 28th ACM international conference on multimedia. 2020. (Year: 2020). [cited by examiner]
Ciftci, Umur Aybars, Ilke Demir, and Lijun Yin. “Fakecatcher: Detection of synthetic portrait videos using biological signals.” IEEE transactions on pattern analysis and machine intelligence (2020 (Year: 2020). [cited by examiner]
Demir, Ilke, and Umur Aybars Ciftci. “Where do deep fakes look? synthetic face detection via gaze tracking.” ACM symposium on eye tracking research and applications. 2021. (Year: 2021). [cited by examiner]
Mantiuk, Radoslaw, Bartosz Bazyluk, and Rafal K. Mantiuk. “Gaze-driven object tracking for real time rendering.” Computer Graphics Forum. vol. 32. No. 2pt2. Oxford, UK: Blackwell Publishing Ltd, 2013. (Year: 2013). [cited by examiner]
Brau, Ernesto, et al. “Multiple-gaze geometry: Inferring novel 3d locations from gazes observed in monocular video.” Proceedings of the European Conference on Computer Vision (ECCV). 2018 (Year: 2018). [cited by examiner]
“Content Authenticity Initiative”, Adobe Inc., [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://contentauthenticity.org/>., 2020, 2 Pages. [cited by applicant]
Afchar, Darius , et al., “MesoNet: a Compact Facial Video Forgery Detection Network”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1809.00888.pdf>., S… [cited by applicant]
Dolhansky, Brian , et al., “The Deepfake Detection Challenge (DFDC) Preview Dataset”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1910.08854.pdf>., O… [cited by applicant]
Dosovitskiy, Alexey , et al., “FlowNet: Learning Optical Flow With Convolutional Networks”, IEEE International Conference on Computer Vision [retrieved Mar. 21, 2022]. Retrieved from the Internet <http://citeseerx.ist.p… [cited by applicant]
Dufour, Nick , et al., “Contributing Data to Deepfake Detection Research”, Google AI Blog [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://ai.googleblog.com/2019/09/contributing-data-to-deepfake-detection… [cited by applicant]
Jiang, Liming , et al., “DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/p… [cited by applicant]
Kleinsmith, Andrea , et al., “Affective Body Expression Perception and Recognition: A Survey”, IEEE Transactions on Affective Computing, vol. 4, No. 1 [retrieved Mar. 21, 2022]. Retrieved from the Internet <http://web4.… [cited by applicant]
Korshunov, Pavel , et al., “DeepFakes: a New Threat to Face Recognition? Assessment and Detection”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1812.… [cited by applicant]
Kosti, Ronak , et al., “EMOTIC: Emotions in Context Dataset”, IEEE Conference on Computer Vision and Pattern Recognition Workshops [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://openaccess.thecvf.com/co… [cited by applicant]
Li, Yuezun , et al., “Celeb-DF: A New Dataset for DeepFake Forensics”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1909.12962v2.pdf>., Oct. 2, 2019, … [cited by applicant]
Li, Yuezun , et al., “Exposing DeepFake Videos by Detecting Face Warping Artifacts”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1811.00656.pdf>., Ma… [cited by applicant]
Mittal, Trisha , et al., “EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's Principle”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pd… [cited by applicant]
Mittal, Trisha , et al., “Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective Cues”, Proceedings of the 28th ACM International Conference on Multimedia [retrieved Mar. 21, 2022]. Retrieved from… [cited by applicant]
Nguyen, Huy H, et al., “Capsule-Forensics: Using Capsule Networks to Detect Forged Images and Videos”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/18… [cited by applicant]
Nguyen, Huy H, et al., “Multi-task Learning for Detecting and Segmenting Manipulated Facial Images and Videos”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.o… [cited by applicant]
Recasens, ADRIà, et al., “Where are they looking?”, PHD thesis, Massachusetts Institute of Technology [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://proceedings.neurips.cc/paper/2015/file/ec8956637a9978… [cited by applicant]
Rossler, Andreas , et al., “FaceForensics++: Learning to Detect Manipulated Facial Images”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <http://128.84.21.203/pdf/1901.08971… [cited by applicant]
Yang, Xin , et al., “Exposing Deep Fakes Using Inconsistent Head Poses”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1811.00661.pdf>., Nov. 13, 2018,… [cited by applicant]
Zhou, Peng , et al., “Two-Stream Neural Networks for Tampered Face Detection”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1803.11276.pdf>., Mar. 29,… [cited by applicant]
Zi, Bojia , et al., “WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection”, Cornell University arXiv, arXiv.org [retrieved Mar. 21, 2022]. Retrieved from the Internet <http://128.84.4.18/pdf/2101.01456>… [cited by applicant]