IP Library › Granted Patent US 12,475,686
Granted Patent B2
US 12,475,686 · App. 19/094,609 · Granted Nov 18, 2025

Common sense reasoning for deepfake detection

Inventors: Gaurav Bharaj (San Francisco, CA); Yue Zhang (East Lansing, MI); Ben Colman (New York, NY); Ali Shahriyari (Las Vegas, NV)
Assignee: Reality Defender, Inc.
G06V10/7715G06V40/168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,686
App. No.
19/094,609
Granted
Nov 18, 2025
Kind
B2
Abstract

An exemplary method for detecting deepfake images and providing customized analysis comprises: receiving, from a user, a textual user inquiry regarding an image; inputting the textual inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises: an image encoder for generating a plurality of image embeddings based on the image; a text encoder for generating a plurality of textual embeddings based on the textual inquiry; one or more layers for generating a plurality of answer embeddings; and a language model for generating a textual analysis based on the plurality of answer embeddings; and outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes one or more visual features in the image and one or more attributes of the one or more visual features that contribute to the classification result.

Claims (39)

1 . A method for detecting deepfake images and providing customized analysis, comprising:

receiving a user inquiry regarding an image, the user inquiry comprising a question about whether one or more visual features in the image are fake;

inputting a text string associated with the user inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating one or more representations of the image;

a text encoder for generating one or more representations of the text string; and

one or more layers for generating a textual analysis based on the one or more representations of the image and the one or more representations of the text string; and

outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes a classification result of whether the one or more visual features in the image are fake and one or more attributes of the one or more visual features that contribute to the classification result.

2 . The method of claim 1 , wherein the visual features include facial features.

3 . The method of claim 2 , wherein the facial features include eyebrows, skin, eyes, nose, mouth, teeth, chin, hair, accessories, shadow, or a combination thereof.

4 . The method of claim 1 , wherein the user inquiry further comprises a question about whether the image is fake.

5 . The method of claim 1 , further comprising: displaying a chatbot user interface for receiving the user inquiry related to the image.

6 . The method of claim 1 , wherein the one or more representations of the image comprise a plurality of image embeddings and wherein the one or more representations of the text string comprise a plurality of textual embeddings.

7 . The method of claim 6 , wherein the one or more layers comprises a set of layers for generating a plurality of answer embeddings based on the plurality of image embeddings and the plurality of textual embeddings.

8 . The method of claim 7 , wherein the set of layers of the deepfake detection model further comprises: a plurality of cross attention layers.

9 . The method of claim 8 , further comprising generating, using the plurality of cross-attention layers, a plurality of encoded image embeddings and a plurality of encoded text embeddings based on the plurality of image embeddings and the plurality of textual embeddings.

10 . The method of claim 9 , wherein the set of layers of the deepfake detection model further comprises: a text decoder.

11 . The method of claim 10 , further comprising generating, using the text decoder, a plurality of decoded text embeddings.

12 . The method of claim 11 , wherein the plurality of answer embeddings comprises the plurality of decoded text embeddings.

13 . The method of claim 10 , wherein the text decoder is trained via text contrastive learning.

14 . The method of claim 1 , wherein the image encoder is trained via image contrastive learning.

15 . The method of claim 1 , wherein the deepfake detection model is trained using a training dataset comprising a training image, a corresponding textual inquiry regarding the training image, a classification result of whether the training image is fake, and a corresponding textual response to the textual inquiry.

16 . The method of claim 15 , wherein the corresponding textual response to the textual inquiry is generated by one or more selections of a plurality of predefined descriptors by a human annotator.

17 . The method of claim 15 , wherein the deepfake detection model is trained at least partially by: inputting the corresponding textual response to the textual inquiry into a plurality of causal self-attention layers.

18 . The method of claim 1 , wherein the deepfake detection model comprises a BLIP model.

19 . A system for detecting deepfake images and providing customized analysis, comprising: one or more processors; one or more memories; and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a user inquiry regarding an image, the user inquiry comprising a question about whether one or more visual features in the image are fake;

inputting a text string associated with the user inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating one or more representations of the image;

a text encoder for generating one or more representations of the text string; and

one or more layers for generating a textual analysis based on the one or more representations of the image and the one or more representations of the text string; and

outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes a classification result of whether the one or more visual features in the image are fake and one or more attributes of the one or more visual features that contribute to the classification result.

20 . A non-transitory computer-readable storage medium storing one or more programs for detecting deepfake images and providing customized analysis, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform:

receive a user inquiry regarding an image, the user inquiry comprising a question about whether one or more visual features in the image are fake;

input a text string associated with the user inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating one or more representations of the image;

a text encoder for generating one or more representations of the text string; and

one or more layers for generating a textual analysis based on the one or more representations of the image and the one or more representations of the text string; and

output the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes a classification result of whether the one or more visual features in the image are fake and one or more attributes of the one or more visual features that contribute to the classification result.

21 . The method of claim 1 , wherein the textual analysis further comprises one or more attributes of one or more additional visual features, wherein the one or more additional features are different from the one or more visual features specified in the user inquiry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2025
From: BHARAJ, GAURAV; ZHANG, YUE; COLMAN, BEN; SHAHRIYARI, ALI
To: REALITY DEFENDER, INC.
Reel/Frame 072163/0329 →
Continuity (3)
Continuation 18751168 · Jun 21, 2024
Provisional Application 63600579 · Nov 17, 2023
Related Publication 20250225773A1 · Jul 10, 2025
References Cited (58)
US 12288379B1 · Bharaj · 2025 [cited by examiner]
US 20220058340A1 · Aggarwal · 2022 [cited by applicant]
US 20230153599A1 · Dalli · 2023 [cited by applicant]
US 20230154188A1 · Li · 2023 [cited by applicant]
US 20230401824A1 · Khan · 2023 [cited by applicant]
US 20240291779A1 · Catalano · 2024 [cited by examiner]
Chang et al., “Antifakeprompt: Prompt-Tuned Visionlanguage Models Are Fake Image Detectors”, Nov. 3, 2023. (Year: 2023). [cited by examiner]
Sun et al., “Towards General Visual-Linguistic Face Forgery Detection”, Jul. 31, 2023. (Year: 2023). [cited by examiner]
Agrawal et al. “nocaps: novel object captioning at scale.” IEEE/CVF International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea; pp. 1-23. [cited by applicant]
Agrawal et al. “VQA: Visual Question Answering,” IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, Santiago, Chile; pp. 1-25. [cited by applicant]
Alayrac et al. “Flamingo: a Visual Language Model for Few-Shot Learning,” NeurIPS 2022: 36th Conference on Neural Information Processing Systems, Nov. 28-Dec. 9, 2022, New Orleans, Louisiana; pp. 1-54. [cited by applicant]
Anderson et al. “SPICE: Semantic Propositional Image CaptionEvaluation,” Computer Vision—ECCV 2016: 14th European Conference, Oct. 11-14, 2016, Amsterdam, The Netherlands; pp. 1-17. [cited by applicant]
Anderson et al. “Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 18-23, 2018, Salt… [cited by applicant]
Bai et al. “AUNet: Learning Relations Between Action Units for Face Forgery Detection,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2023, Vancouver, BC, Canada; pp. 24709-24719. [cited by applicant]
Bharaj et al., U.S. Office Action dated Nov. 7, 2024, directed to U.S. Appl. No. 18/751,168; 11 pages. [cited by applicant]
C. Lin. “ROUGE: A Package for Automatic Evaluation of Summaries,” Workshop on Text Summarization Branches Out, Jul. 25-26, 2004, Barcelona, Spain; 24 pages. [cited by applicant]
Chang et al. (Nov. 2023). “AntifakePrompt: Prompt-Tuned Vision-Language Models are Fake Image Detectors,” located at https://arxiv.org/abs/2310.17419; pp. 1-24. [cited by applicant]
Coccomini et al. “Combining EfficientNet and Vision Transformers for Video Deepfake Detection,” International Conference on Image Analysis and Processing, May 23-27, 2022, Lecce, Italy; pp. 1-11. [cited by applicant]
Dai et al. (Jun. 2023). “InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning,” located at https://arxiv.org/abs/2305.06500; pp. 1-17. [cited by applicant]
Denkowski et al. “Meteor Universal: Language Specific Translation Evaluation for Any Target Language,” Ninth Workshop on Statistical Machine Translation, Jun. 26-27, 2014, Baltimore, Maryland; pp. 376-380. [cited by applicant]
Devlin et al. (Oct. 2018). “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” located at https://arxiv.org/abs/1810.04805; 16 pages. [cited by applicant]
Dolhansky et al. (Oct. 2020). “The DeepFake Detection Challenge (DFDC) Dataset,” located at https://arxiv.org/abs/2006.07397; pp. 1-13. [cited by applicant]
Dosovitskiy et al. (Oct. 2020). “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” located at https://arxiv.org/abs/2010.11929; pp. 1-22. [cited by applicant]
Draelos et al. (Nov. 2020). “Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks,” located at https://arxiv.org/abs/2011.08891; pp. 1-20. [cited by applicant]
F. Chollet. “Xception: Deep Learning with Depthwise Separable Convolutions,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, Honolulu, Hawaii; pp. 1251-1258. [cited by applicant]
Goodfellow et al. (Jun. 2014). “Generative Adversarial Nets,” Communications of the ACM 63(11); pp. 1-9. [cited by applicant]
Goyal et al. “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, Honolulu, Hawaii; p… [cited by applicant]
Guo et al. “Controllable Guide-Space for Generalizable Face Forgery Detection,” IEEE/CVF International Conference on Computer Vision, Oct. 1-6, 2023, Paris, France; pp. 1-13. [cited by applicant]
Gupta et al. “Visual Programming: Compositional visual reasoning without training,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 17-24, 2023, Vancouver, BC, Canada; 13 pages. [cited by applicant]
Ho et al. (Jun. 2020) “Denoising Diffusion Probabilistic Models,” Advances in neural information processing systems 33; pp. 1-25. [cited by applicant]
Hong et al. “VLN BERT: A Recurrent Vision-and-Language BERT for Navigation,” IEEE/CVF conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, Nashville, Tennessee; pp. 1-18. [cited by applicant]
Karras et al. “Analyzing and Improving the Image Quality of StyleGAN,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; pp. 1-21. [cited by applicant]
Lake et al. (2017). “Building Machines That Learn and Think Like People,” Behavioral and Brain Sciences 40; pp. 1-58. [cited by applicant]
Li et al. “Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; 10 pages. [cited by applicant]
Li et al. “Face X-ray for More General Face Forgery Detection,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; 10 pages. [cited by applicant]
Li et al. (Jun. 2023). “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” located at https://arxiv.org/abs/2301.12597; 13 pages. [cited by applicant]
Mirsky et al. (Jan. 2020). “The Creation and Detection of Deepfakes: A Survey,” ACM Computing Surveys 1(1); pp. 1:1-1:38. [cited by applicant]
Oord et al. (Jul. 2018). “Representation Learning with Contrastive Predictive Coding,” located at https://arxiv.org/abs/1807.03748; pp. 1-13. [cited by applicant]
Papineni et al. “BLEU: a Method for Automatic Evaluation of Machine Translation,” 40th Annual Meeting on Association for Computational Linguistics, Jul. 7-12, 2002, Philadelphia, Pennsylvania; 8 pages. [cited by applicant]
Ramesh et al. (Apr. 2022). “Hierarchical Text-Conditional Image Generation with CLIP Latents,” located at https://arxiv.org/abs/2204.06125; pp. 1-27. [cited by applicant]
Ricker et al. (Oct. 2022). “Towards the Detection of Diffusion Model Deepfakes,” located at https://arxiv.org/abs/2210.14571; 31 pages. [cited by applicant]
Rossler et al. “FaceForensics++: Learning to Detect Manipulated Facial Images,” IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Oct. 27-28, 2019, Seoul, Korea; pp. 1-11. [cited by applicant]
Saharia et al. (May 2022). “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” Advances in Neural Information Processing Systems 35; pp. 1-46. [cited by applicant]
Schwenk et al. “A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge,” European Conference on Computer Vision, Oct. 23-27, 2022, Tel Aviv, Israel; pp. 1-20. [cited by applicant]
Selvaraju et al. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy; pp. 1-23. [cited by applicant]
Shiohara et al. “Detecting Deepfakes with Self-Blended Images,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, New Orleans, Louisiana; pp. 1-11. [cited by applicant]
Sun et al. (Jul. 2023). “Towards General Visual-Linguistic Face Forgery Detection,” located at https://arxiv.org/abs/2307.16545; 16 pages. [cited by applicant]
T. Geller, (Jul. 2008). “Overcoming the Uncanny Valley,” IEEE Computer Graphics and Applications 28(4); pp. 11-17. [cited by applicant]
Tan et al. “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” International Conference on Machine Learning, Jun. 9-15, 2019, Long Beach, California; 11 pages. [cited by applicant]
Tolosana et al. (Jan. 2020). “DeepFakes and Beyond: A Survey of Face Manipulation and Fake Detection,” Information Fusion 64; pp. 1-23. [cited by applicant]
Trinh et al. (Jun. 2018). “A Simple Method for Commonsense Reasoning,” located at https://arxiv.org/abs/1806.02847; pp. 1-12. [cited by applicant]
Turton et al. (Jan. 2020). “How Deepfakes Make Disinformation More Real Than Ever: QuickTake,” Bloomberg Industry Group, Inc.; pp. 1-3. [cited by applicant]
Vaccari et al. (Jan. 2020). “Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News,” Social Media+ Society 6(1); pp. 1-13. [cited by applicant]
Vedantam et al. “CIDEr: Consensus-based Image Description Evaluation,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 7-12, 2015, Boston, Massachusetts; pp. 1-17. [cited by applicant]
Vinyals et al. (Sep. 2016). “Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge,” IEEE Transaction on Pattern Analysis and Machine Intelligence; pp. 1-12. [cited by applicant]
Zhu et al. (Apr. 2023). “MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,” located at https://arxiv.org/abs/2304.10592; pp. 1-15. [cited by applicant]
Zhuang et al. “UIA-ViT: Unsupervised Inconsistency-Aware Method based on Vision Transformer for Face Forgery Detection,” European Conference on Computer Vision (ECCV 2022), Oct. 23-27, 2022, Tel Aviv, Israel; pp. 1-16. [cited by applicant]
International Search Report and Written Opinion mailed Dec. 27, 2024, directed to International Application No. PCT/US2024/056145; 9 pages. [cited by applicant]