IP Library › Granted Patent US 12,462,813
Granted Patent B1
US 12,462,813 · App. 19/200,316 · Granted Nov 4, 2025

Data-driven audio deepfake detection

Inventors: Gaurav Bharaj (Los Angeles, CA); Petr Grinberg (Lausanne, CH); Ankur Kumar (Los Angeles, CA); Surya Koppisetti (Coquitlam, CA)
Assignee: Reality Defender, Inc.
G10L17/26G10L17/18G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,813
App. No.
19/200,316
Granted
Nov 4, 2025
Kind
B1
Abstract

An exemplary method for generating a visual representation of manipulations in an audio signal includes inputting the audio signal into a trained machine-learning model, wherein the machine-learning model is trained by generating, based on a training bona fide audio signal, a training bona fide time-frequency representation; generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal; generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and training the audio deepfake detection machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal; and generating, by the machine-learning model, the visual representation of the manipulations in the audio signal.

Claims (44)

1 . A method for generating a visual representation of manipulations in an audio signal, comprising:

inputting the audio signal into a trained machine-learning model, wherein the machine-learning model is trained by:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the audio deepfake detection machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal; and

generating, by the machine-learning model, the visual representation of the manipulations in the audio signal.

2 . The method of claim 1 , further comprising displaying the visual representation of the manipulations in the audio signal.

3 . The method of claim 1 , further comprising displaying one or more timestamps associated with the manipulations in the audio signal.

4 . The method of claim 1 , further comprising:

inputting the visual representation of the manipulations in the audio signal into a language model; and

generating, by the language model, a natural language description of the manipulations in the audio signal.

5 . The method of claim 4 , wherein the language model is trained based on the training visual representation of the manipulations in the training spoofed audio signal and a natural-language description of the manipulations in the training spoofed audio signal.

6 . The method of claim 4 , further comprising displaying the natural language description of the manipulations in the audio signal.

7 . The method of claim 1 , wherein the audio signal comprises a real-time signal from a telephone call.

8 . The method of claim 7 , further comprising automatically terminating the telephone call.

9 . The method of claim 7 , further comprising displaying a warning that the telephone call comprises manipulations.

10 . The method of claim 7 , further comprising outputting a recommendation to terminate the telephone call.

11 . The method of claim 1 , wherein the audio signal comprises a recorded voicemail message.

12 . The method of claim 11 , further comprising displaying a warning that the recorded voicemail message comprises manipulations.

13 . The method of claim 1 , wherein the training bona fide time-frequency representation and the training spoofed time-frequency representation are spectrograms.

14 . The method of claim 1 , wherein the training bona fide audio signal comprises one or more utterances by a speaker and the training spoofed audio signal comprises a manipulated version of the one or more utterances by the speaker.

15 . The method of claim 1 , wherein the training spoofed audio signal is generated by a deepfake machine-learning model based on the training bona fide audio signal.

16 . The method of claim 1 , wherein the training bona fide audio signal and the training spoofed audio signal are time synchronized.

17 . The method of claim 1 , wherein generating the training visual representation of the manipulations in the training spoofed audio signal further comprises:

smoothing the training bona fide time-frequency representation; and

smoothing the training spoofed time-frequency representation, wherein the difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation is calculated by subtracting the smoothed training bona fide time-frequency representation from the smoothed training spoofed time-frequency representation.

18 . The method of claim 1 , wherein generating the training visual representation of the manipulations in the training spoofed audio signal further comprises normalizing the difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation.

19 . The method of claim 1 , wherein the machine-learning model comprises a diffusion model.

20 . The method of claim 1 , wherein the machine-learning model comprises a convolutional neural network.

21 . A system for generating a visual representation of manipulations in an audio signal, comprising: one or more processors; one or more memories; and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:

inputting the audio signal into a trained machine-learning model, wherein the machine-learning model is trained by:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the audio deepfake detection machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal; and

generating, by the machine-learning model, the visual representation of the manipulations in the audio signal.

22 . A non-transitory computer-readable storage medium storing one or more programs for generating a visual representation of manipulations in an audio signal, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform:

inputting the audio signal into a trained machine-learning model, wherein the machine-learning model is trained by:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the audio deepfake detection machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal; and

generating, by the machine-learning model, the visual representation of the manipulations in the audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2025
From: BHARAJ, GAURAV; GRINBERG, PETR; KUMAR, ANKUR; KOPPISETTI, SURYA
To: REALITY DEFENDER, INC.
Reel/Frame 072466/0635 →
Continuity (1)
Provisional Application 63760008 · Feb 18, 2025
References Cited (52)
US 11929078B2 · Tuo · 2024 [cited by examiner]
US 12189712B1 · Bharaj · 2025 [cited by examiner]
US 12288379B1 · Bharaj · 2025 [cited by examiner]
US 20200035247A1 · Boyadjiev · 2020 [cited by examiner]
US 20230274758A1 · Markhasin · 2023 [cited by examiner]
US 20230360653A1 · Yan · 2023 [cited by examiner]
US 20240169042A1 · Kong · 2024 [cited by examiner]
US 20250094837A1 · Malik · 2025 [cited by examiner]
US 20250148788A1 · Bellinger · 2025 [cited by examiner]
US 20250166358A1 · Bharaj · 2025 [cited by examiner]
US 20250182510A1 · Malik · 2025 [cited by examiner]
Achtibat et al. “AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers,” Proceedings of the 41st International Conference on Machine Learning, Jul. 21-Jul. 27, 2024, Vienna, Austria; pp. 1-34. [cited by applicant]
Amit et al. (2021). “SegDiff: Image Segmentation with Diffusion Probabilistic Models,” arXiv:2112.00390: 1-13. [cited by applicant]
Arrieta et al. (2020). “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Information Fusion 58: 82-115. [cited by applicant]
Bai et al. (Jul. 2024). “Diffusion-Based Adversarial Purification for Speaker Verification,” IEEE Signal Processing Letters, Journal of Latex Class Filed, 14(8): 1-5. [cited by applicant]
Cazenavette et al. “FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable Diffusion,” Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. … [cited by applicant]
Chefer et al. “Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 11-17, 2021, Online, Canad… [cited by applicant]
Cordts et al. “The Cityscapes Dataset for Semantic Urban Scene Understanding,” Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27-30, 2016, Las Vegas, USA; pp. 3213-3223. [cited by applicant]
Frank et al. “WaveFake: A Data Set to Facilitate Audio Deepfake Detection,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks, Dec. 6-14, 2001, Virtual-only Confere… [cited by applicant]
Ge et al. “Explainable deepfake and spoofing detection: an attack analysis using Shapley Additive explanations,” The Speaker and Language Recognition Workshop (Odyssey 2022), Jun. 28-Jul. 1, 2022, Beijing, China; pp. 70… [cited by applicant]
Ge et al. “Explaining Deep Learning Models for Spoofing and Deepfake Detection with Shapley Additive Explanations,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 23-27, 2022, Si… [cited by applicant]
Grinberg et al. (Jan. 2025). “What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain,” arXiv:2501.13887: 1-5. [cited by applicant]
Ho et al. (Jun. 2020) “Denoising Diffusion Probabilistic Models,” Advances in neural information processing systems 33; pp. 1-25. [cited by applicant]
Jadon “A survey of loss functions for semantic segmentation,” 2020 IEEE conference on computational intelligence in bioinformatics and computational biology (CIBCB), Oct. 27-29, 2020, Vina del Mar, Chile; 6 pages. [cited by applicant]
Kim et al. “Diff-SV: A Unified Hierarchical Framework For Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic Models,” 2024 IEEE International Conference on Acoustics, Speech and Signal Processin… [cited by applicant]
Kumar et al. (Mar. 2017). “A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology,” IEEE Transactions on Medical Imaging 36(7): 1-11. [cited by applicant]
Kwak et al. (May 2023). “Voice Spoofing Detection Through Residual Network, Max Feature Map, and Depthwise Separable Convolution,” IEEE Access 11: 49140-49152. [cited by applicant]
Li et al. (Apr. 2024). “Audio Anti-Spoofing Detection: A Survey,” arXiv:2404.13914: 1-43. [cited by applicant]
Li et al. (Jun. 2024). “Interpretable Temporal Class Activation Representation for Audio Spoofing Detection,” Interspeech 2024, arXiv:2406.08825: 5 pages. [cited by applicant]
Lim et al. (Apr. 13, 2022). “Detecting Deepfake Voice Using Explainable Deep Learning Techniques,” Applied Sciences 12(3926): 1-15. [cited by applicant]
Liu et al. (Jun. 2024). “How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?,” Interspeech 2024, arXiv:2406.02483: 5 pages. [cited by applicant]
Mancini et al. (Sep. 2024). “LMAC-TD: Producing Time Domain Explanations for Audio Classifiers,” arXiv:2409.08655: 1-5. [cited by applicant]
Müller et al. “Does Audio Deepfake Detection Generalize?,” Proceedings of the 2022 Interspeech Conference, Sep. 18-22, 2022, Incheon, Korea; 5 pages. [cited by applicant]
Paissan et al. “Listenable Maps for Audio Classifiers,” Proceedings of the 41st International Conference on Machine Learning, Jul. 21-Jul. 27, 2024, Vienna, Austria; pp. 1-13. [cited by applicant]
Prenger et al. “Waveglow: A Flow-Based Generative Network for Speech Synthesis,” ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 12-17, 2019, Brighton, United King… [cited by applicant]
Rottensteiner et al. (2014). “ISPRS Semantic Labeling Contest, ISPRS: Leopoldsh”ohe, Germany 1(4): 1 page. [cited by applicant]
Salvi et al. “Towards Frequency Band Explainability in Synthetic Speech Detection,” 31st European Signal Processing Conference (EUSIPCO), Sep. 4-8, 2023, Helsinki, Finland; pp. 620-624. [cited by applicant]
Selvaraju et al. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” IEEE International Conference on Computer Vision (ICCV), Oct. 22-29, 2017, Venice, Italy; pp. 1-23. [cited by applicant]
Sudre et al. “Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations,” Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: Third Internat… [cited by applicant]
Sun et al. “AI-Synthesized Voice Detection Using Neural Vocoder Artifacts,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 17-24, 2023, Vancouver, BC, Canada; pp. 904-9… [cited by applicant]
Sun et al. “DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion,” 38th Conference on Neural Information Processing Systems (NeurIPS 2024), Dec. 9-15, 2024, Vancouver, BC, Canada; 14… [cited by applicant]
Tak et al. “An explainability study of the constant Q cepstral coefficient spoofing countermeasure for automatic speaker verification,” The Speaker and Language Recognition Workshop (Odyssey 2020), Nov. 1-5, 2020, Tokyo… [cited by applicant]
Tak et al. “Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation,” The Speaker and Language Recognition Workshop (Odyssey 2022), Jun. 28-Jul. 1, 2022, Beijing, China; 10… [cited by applicant]
Tak et al. “Rawboost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing,” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICAS… [cited by applicant]
Wang et al. (2020). “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” Computer Speech & Language, 64: 1-24. [cited by applicant]
Wang et al. (Apr. 2004). “Image Quality Assessment: From Error Visibility to Structural Similarity,” IEEE Transactions on Image Processing 13(4): 1-14. [cited by applicant]
Wang et al. (Nov. 2019). “Neural source-filter waveform models for statistical parametric speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing 28: pp. 1-14. [cited by applicant]
Wang et al. (Sep. 2018) “ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks,” located at https://doi.org/10.48550/arXiv, 1809.00219, (23 pages). [cited by applicant]
Wang et al. “Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?,” ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICA… [cited by applicant]
Wang et al. “Spoofed Training Data for Speech Spoofing Countermeasure Can Be Efficiently Created Using Neural Vocoders,” ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)… [cited by applicant]
Zhang et al. (Sep. 2023). “The Impact of Silence on Speech Anti-Spoofing,” IEEE/ACM Transactions on Audio, Speech, and Language Processing 31: 3374-3389. [cited by applicant]
Zhang et al. “Common Sense Reasoning for Deepfake Detection,” 18th European Conference on Computer Vision, Sep. 29-Oct. 4, 2024, Milan, Italy; pp. 1-27. [cited by applicant]