IP Library › Granted Patent US 12,288,379
Granted Patent B1
US 12,288,379 · App. 18/751,168 · Granted Apr 29, 2025

Common sense reasoning for deepfake detection

Inventors: Gaurav Bharaj (San Francisco, CA); Yue Zhang (East Lansing, MI); Ben Colman (New York, NY); Ali Shahriyari (Las Vegas, NV)
Assignee: Reality Defender, Inc.
G06V10/7715G06V40/168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,379
App. No.
18/751,168
Granted
Apr 29, 2025
Kind
B1
Abstract

An exemplary method for detecting deepfake images and providing customized analysis comprises: receiving, from a user, a textual user inquiry regarding an image; inputting the textual inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises: an image encoder for generating a plurality of image embeddings based on the image; a text encoder for generating a plurality of textual embeddings based on the textual inquiry; one or more layers for generating a plurality of answer embeddings; and a language model for generating a textual analysis based on the plurality of answer embeddings; and outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes one or more visual features in the image and one or more attributes of the one or more visual features that contribute to the classification result.

Claims (40)

1. A method for detecting deepfake images and providing customized analysis, comprising:

receiving, from a user, a textual user inquiry regarding an image;

inputting the textual inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating a plurality of image embeddings based on the image;

a text encoder for generating a plurality of textual embeddings based on the textual inquiry;

one or more layers for generating a plurality of answer embeddings based on the plurality of image embeddings and the plurality of textual embeddings; and

a language model for generating a textual analysis based on the plurality of answer embeddings; and

outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes one or more visual features in the image and one or more attributes of the one or more visual features that contribute to the classification result.

2. The method of claim 1 , wherein the visual features include facial features.

3. The method of claim 2 , wherein the facial features include eyebrows, skin, eyes, nose, mouth, teeth, chin, hair, accessories, shadow, or a combination thereof.

4. The method of claim 1 , wherein the textual user inquiry comprises a question about whether the image is fake.

5. The method of claim 1 , wherein the textual user inquiry comprises a question about whether one or more visual feature in the image are fake.

6. The method of claim 1 , further comprising: displaying a chatbot user interface for receiving one or more textual inquiries related to the image.

7. The method of claim 1 , wherein the one or more layers of the deepfake detection model further comprises: a plurality of cross attention layers.

8. The method of claim 7 , further comprising generating, using the plurality of cross-attention layers, a plurality of encoded image embeddings and a plurality of encoded text embeddings based on the plurality of image embeddings and the plurality of text embeddings.

9. The method of claim 8 , wherein the one or more layers of the deepfake detection model further comprises: a text decoder.

10. The method of claim 9 , further comprising generating, using the text decoder, a plurality of decoded text embeddings.

11. The method of claim 10 , wherein the plurality of answer embeddings comprises the plurality of decoded text embeddings.

12. The method of claim 9 , wherein the text decoder is trained via text contrastive learning.

13. The method of claim 1 , wherein the image encoder is trained via image contrastive learning.

14. The method of claim 1 , wherein the deepfake detection model is trained using a training dataset comprising a training image, a corresponding textual inquiry regarding the training image, a classification result of whether the training image is fake, and a corresponding textual response to the textual inquiry.

15. The method of claim 14 , wherein the corresponding textual response to the textual inquiry is generated by one or more selections of a plurality of predefined descriptors by a human annotator.

16. The method of claim 14 , wherein the deepfake detection model is trained at least partially by: inputting the corresponding textual response to the textual inquiry into a plurality of causal self-attention layers.

17. The method of claim 1 , wherein the deepfake detection model comprises a BLIP model.

18. A system for detecting deepfake images and providing customized analysis, comprising: one or more processors; one or more memories; and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving, from a user, a textual user inquiry regarding an image;

inputting the textual inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating a plurality of image embeddings based on the image;

a text encoder for generating a plurality of textual embeddings based on the textual inquiry;

one or more layers for generating a plurality of answer embeddings based on the plurality of image embeddings and the plurality of textual embeddings; and

a language model for generating a textual analysis based on the plurality of answer embeddings; and

outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes one or more visual features in the image and one or more attributes of the one or more visual features that contribute to the classification result.

19. A non-transitory computer-readable storage medium storing one or more programs for detecting deepfake images and providing customized analysis, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform:

receiving, from a user, a textual user inquiry regarding an image;

inputting the textual inquiry and the image into a deepfake detection model, wherein the deepfake detection model comprises:

an image encoder for generating a plurality of image embeddings based on the image;

a text encoder for generating a plurality of textual embeddings based on the textual inquiry;

one or more layers for generating a plurality of answer embeddings based on the plurality of image embeddings and the plurality of textual embeddings; and

a language model for generating a textual analysis based on the plurality of answer embeddings; and

outputting the textual analysis, wherein the textual analysis includes a classification result of whether the image is fake and further includes one or more visual features in the image and one or more attributes of the one or more visual features that contribute to the classification result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2025
From: BHARAJ, GAURAV; ZHANG, YUE; COLMAN, BEN; SHAHRIYARI, ALI
To: REALITY DEFENDER, INC.
Reel/Frame 070562/0023 →
Continuity (1)
Provisional Application 63600579 · Nov 17, 2023
References Cited (52)
US 20220058340A1 · Aggarwal · 2022 [cited by examiner]
US 20230153599A1 · Dalli · 2023 [cited by examiner]
US 20230154188A1 · Li · 2023 [cited by examiner]
US 20230401824A1 · Khan · 2023 [cited by examiner]
Chang et al., “ANTIFAKEPROMPT: Prompt-Tuned VISIONLANGUAGE Models Are Fake Image Detectors”, Nov. 3, 2023. (Year: 2023). [cited by examiner]
Dai et al., “InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning”, Jun. 15, 2023. (Year: 2023). [cited by examiner]
Li et al., “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, Jun. 15, 2023. (Year: 2023). [cited by examiner]
Agrawal et al. “nocaps: novel object captioning at scale.” IEEE/CVF International Conference on Computer Vision, Oct. 27-Nov. 2, 2019, Seoul, Korea; pp. 1-23. [cited by applicant]
Agrawal et al. “VQA: Visual Question Answering,” IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, Santiago, Chile; pp. 1-25. [cited by applicant]
Alayrac et al. “Flamingo: a Visual Language Model for Few-Shot Learning,” NeurIPS 2022: 36th Conference on Neural Information Processing Systems, Nov. 28-Dec. 9, 2022, New Orleans, Louisiana; pp. 1-54. [cited by applicant]
Anderson et al. “SPICE: Semantic Propositional Image CaptionEvaluation,” Computer Vision—ECCV 2016: 14th European Conference, Oct. 11-14, 2016, Amsterdam, The Netherlands; pp. 1-17. [cited by applicant]
Anderson et al. “Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Jun. 18-23, 2018, Salt… [cited by applicant]
Bai et al. “AUNet: Learning Relations Between Action Units for Face Forgery Detection,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2023, Vancouver, BC, Canada; pp. 24709-24719. [cited by applicant]
C. Lin. “Rouge: A Package for Automatic Evaluation of Summaries,” Workshop on Text Summarization Branches Out, Jul. 25-26, 2004, Barcelona, Spain; 24 pages. [cited by applicant]
Coccomini et al. “Combining EfficientNet and Vision Transformers for Video Deepfake Detection,” International Conference on Image Analysis and Processing, May 23-27, 2022, Lecce, Italy; pp. 1-11. [cited by applicant]
Denkowski et al. “Meteor Universal: Language Specific Translation Evaluation for Any Target Language,” Ninth Workshop on Statistical Machine Translation, Jun. 26-27, 2014, Baltimore, Maryland; pp. 376-380. [cited by applicant]
Devlin et al. (Oct. 2018). “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” located at https://arxiv.org/abs/1810.04805; 16 pages. [cited by applicant]
Dolhansky et al. (Oct. 2020). “The DeepFake Detection Challenge (DFDC) Dataset,” located at https://arxiv.org/abs/2006.07397; pp. 1-13. [cited by applicant]
Dosovitskiy et al. (Oct. 2020). “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” located at https://arxiv.org/abs/2010.11929; pp. 1-22. [cited by applicant]
Draelos et al. (Nov. 2020). “Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks,” located at https://arxiv.org/abs/2011.08891; pp. 1-20. [cited by applicant]
F. Chollet. “Xception: Deep Learning with Depthwise Separable Convolutions,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, Honolulu, Hawaii; pp. 1251-1258. [cited by applicant]
Goodfellow et al. (Jun. 2014). “Generative Adversarial Nets,” Communications of the ACM 63(11); pp. 1-9. [cited by applicant]
Goyal et al. “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 21-26, 2017, Honolulu, Hawaii; p… [cited by applicant]
Guo et al. “Controllable Guide-Space for Generalizable Face Forgery Detection,” IEEE/CVF International Conference on Computer Vision, Oct. 1-6, 2023, Paris, France; pp. 1-13. [cited by applicant]
Gupta et al. “Visual Programming: Compositional visual reasoning without training,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 17-24, 2023, Vancouver, BC, Canada; 13 pages. [cited by applicant]
Ho et al. (Jun. 2020) “Denoising Diffusion Probabilistic Models,” Advances in neural information processing systems 33; pp. 1-25. [cited by applicant]
Hong et al. “VLN BERT: A Recurrent Vision-and-Language BERT for Navigation,” IEEE/CVF conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, Nashville, Tennessee; pp. 1-18. [cited by applicant]
Karras et al. “Analyzing and Improving the Image Quality of StyleGAN,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; pp. 1-21. [cited by applicant]
Lake et al. (2017).“Building Machines That Learn and Think Like People,” Behavioral and Brain Sciences 40; pp. 1-58. [cited by applicant]
Li et al. “Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; 10 pages. [cited by applicant]
Li et al. “Face X-ray for More General Face Forgery Detection,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, Seattle, Washington; 10 pages. [cited by applicant]
Mirsky et al. (Jan. 2020). “The Creation and Detection of Deepfakes: A Survey,” ACM Computing Surveys 1(1); pp. 1:1-1:38. [cited by applicant]
Oord et al. (Jul. 2018). “Representation Learning with Contrastive Predictive Coding,” located at https://arxiv.org/abs/1807.03748; pp. 1-13. [cited by applicant]
Papineni et al. “BLEU: a Method for Automatic Evaluation of Machine Translation,” 40th Annual Meeting on Association for Computational Linguistics, Jul. 7-12, 2002, Philadelphia, Pennsylvania; 8 pages. [cited by applicant]
Ramesh et al. (Apr. 2022). “Hierarchical Text-Conditional Image Generation with CLIP Latents,” located at https://arxiv.org/abs/2204.06125; pp. 1-27. [cited by applicant]
Ricker et al. (Oct. 2022). “Towards the Detection of Diffusion Model Deepfakes,” located at https://arxiv.org/abs/2210.14571; 31 pages. [cited by applicant]
Rossler et al. “FaceForensics++: Learning to Detect Manipulated Facial Images,” IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Oct. 27-28, 2019, Seoul, Korea; pp. 1-11. [cited by applicant]
Saharia et al. (May 2022). “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” Advances in Neural Information Processing Systems 35; pp. 1-46. [cited by applicant]
Schwenk et al. “A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge,” European Conference on Computer Vision, Oct. 23-27, 2022, Tel Aviv, Israel; pp. 1-20. [cited by applicant]
Selvaraju et al. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy; pp. 1-23. [cited by applicant]
Shiohara et al. “Detecting Deepfakes with Self-Blended Images,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-24, 2022, New Orleans, Louisiana; pp. 1-11. [cited by applicant]
Sun et al. (Jul. 2023). “Towards General Visual-Linguistic Face Forgery Detection,” located at https://arxiv.org/abs/2307.16545; 16 pages. [cited by applicant]
T. Geller, (Jul. 2008). “Overcoming the Uncanny Valley,” IEEE Computer Graphics and Applications 28(4); pp. 11-17. [cited by applicant]
Tan et al. “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” International Conference on Machine Learning, Jun. 9-15, 2019, Long Beach, California; 11 pages. [cited by applicant]
Tolosana et al. (Jan. 2020). “DeepFakes and Beyond: A Survey of Face Manipulation and Fake Detection,” Information Fusion 64; pp. 1-23. [cited by applicant]
Trinh et al. (Jun. 2018). “A Simple Method for Commonsense Reasoning,” located at https://arxiv.org/abs/1806.02847; pp. 1-12. [cited by applicant]
Turton et al. (Jan. 2020). “How Deepfakes Make Disinformation More Real Than Ever: QuickTake,” Bloomberg Industry Group, Inc.; pp. 1-3. [cited by applicant]
Vaccari et al. (Jan. 2020). “Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News,” Social Media+ Society 6(1); pp. 1-13. [cited by applicant]
Vedantam et al. “CIDEr: Consensus-based Image Description Evaluation,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 7-12, 2015, Boston, Massachusetts; pp. 1-17. [cited by applicant]
Vinyals et al. (Sep. 2016). “Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge,” IEEE Transaction on Pattern Analysis and Machine Intelligence; pp. 1-12. [cited by applicant]
Zhu et al. (Apr. 2023). “MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,” located at https://arxiv.org/abs/2304.10592; pp. 1-15. [cited by applicant]
Zhuang et al. “UIA-ViT: Unsupervised Inconsistency-Aware Method based on Vision Transformer for Face Forgery Detection,” European Conference on Computer Vision (ECCV 2022), Oct. 23-27, 2022, Tel Aviv, Israel; pp. 1-16. [cited by applicant]
Cited By (3)
US 12,462,813 US 12,475,686 US 12,586,395