IP Library › Patent Application 19352136
Patent Application
App. No. 19/352,136

DATA-DRIVEN AUDIO DEEPFAKE DETECTION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/352,136
Abstract

An exemplary method for generating a visual representation of manipulations in an audio signal includes inputting the audio signal into a trained machine-learning model, wherein the machine-learning model is trained by generating, based on a training bona fide audio signal, a training bona fide time-frequency representation; generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal; generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and training the audio deepfake detection machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal; and generating, by the machine-learning model, the visual representation of the manipulations in the audio signal.

Claims (35)

1 . A method for training a machine-learning model to generate a visual representation of manipulations in an audio signal, comprising:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal.

2 . The method of claim 1 , wherein the training bona fide time-frequency representation and the training spoofed time-frequency representation are spectrograms.

3 . The method of claim 1 , wherein the training visual representation of manipulations in the training spoofed audio signal comprises one or more annotations indicating time or frequency information regarding the manipulations in the training spoofed audio signal.

4 . The method of claim 1 , wherein the training bona fide audio signal comprises one or more utterances by a speaker and the training spoofed audio signal comprises a manipulated version of the one or more utterances by the speaker.

5 . The method of claim 1 , wherein the training spoofed audio signal is generated by a deepfake machine-learning model based on the training bona fide audio signal.

6 . The method of claim 4 , wherein the training spoofed audio signal is generated by at least one of a text-to-speech system, a voice cloner, an autoencoder, a speech synthesis model, and a vocoder.

7 . The method of claim 1 , wherein the training bona fide audio signal and the training spoofed audio signal are time synchronized.

8 . The method of claim 7 , wherein the training bona fide audio signal and the training spoofed audio signal are time synchronized using dynamic time warping prior to generating the training bona fide time-frequency representation and training spoofed time-frequency representation.

9 . The method of claim 1 , wherein generating the training visual representation of the manipulations in the training spoofed audio signal further comprises:

smoothing the training bona fide time-frequency representation; and

smoothing the training spoofed time-frequency representation, wherein the difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation is calculated by subtracting the smoothed training bona fide time-frequency representation from the smoothed training spoofed time-frequency representation.

10 . The method of claim 1 , wherein generating the training visual representation of the manipulations in the training spoofed audio signal further comprises normalizing the difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation.

11 . The method of claim 10 , wherein generating the training visual representation of manipulations in the training spoofed audio signal further comprises: applying a threshold to the normalized difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation to produce a binarized mask indicating manipulated regions.

12 . The method of claim 1 , wherein the machine-learning model is trained to provide classifier-specific visual explanations and the method further comprises conditioning training of the machine-learning model on intermediate features extracted from a pre-trained audio deepfake detection classifier.

13 . The method of claim 1 , wherein the machine-learning model is trained to provide classifier-agnostic visual explanations of manipulations in audio signals.

14 . The method of claim 1 , further comprising augmenting the training bona fide audio signal and the training spoofed audio signal with a same random noise parameter prior to generating the training bona fide time-frequency representation and training spoofed time-frequency representation.

15 . The method of claim 1 , wherein the machine-learning model comprises a diffusion model.

16 . The method of claim 1 , wherein the machine-learning model comprises a convolutional neural network.

17 . A system for training a machine-learning model to generate a visual representation of manipulations in an audio signal, comprising: one or more processors; one or more memories;

and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal.

18 . The system of claim 17 , wherein the training bona fide audio signal comprises one or more utterances by a speaker and the training spoofed audio signal comprises a manipulated version of the one or more utterances by the speaker.

19 . A non-transitory computer-readable storage medium storing one or more programs for training a machine-learning model to generate a visual representation of manipulations in an audio signal, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform:

generating, based on a training bona fide audio signal, a training bona fide time-frequency representation;

generating, based on a training spoofed audio signal, a training spoofed time-frequency representation, wherein the training spoofed audio signal is a manipulated version of the training bona fide audio signal;

generating a training visual representation of manipulations in the training spoofed audio signal based at least on a difference between the training bona fide time-frequency representation and the training spoofed time-frequency representation; and

training the machine-learning model based on the training visual representation of the manipulations in the training spoofed audio signal.

20 . The non-transitory computer-readable storage medium of claim 19 , wherein the training bona fide audio signal comprises one or more utterances by a speaker and the training spoofed audio signal comprises a manipulated version of the one or more utterances by the speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2026
From: BHARAJ, GAURAV; GRINBERG, PETR; KUMAR, ANKUR; KOPPISETTI, SURYA
To: REALITY DEFENDER, INC.
Reel/Frame 075448/0052 →