IP Library Patent Application 18777278
Patent Application
App. No. 18/777,278

CENTRALIZED SYNTHETIC SPEECH DETECTION SYSTEM USING WATERMARKING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/777,278
Abstract

Disclosed are systems and methods including software processes executed by a server for obtaining, by a computer, an audio signal including synthetic speech, extracting, by the computer, metadata from a watermark of the audio signal by applying a set of keys associated with a plurality of text-to-speech (TTS) services to the audio signal, the metadata indicating an origin of the synthetic speech in the audio signal, and generating, by the computer, based on the extracted metadata, a notification indicating that the audio signal includes the synthetic speech.

Claims (39)

1 . A computer-implemented method comprising:

obtaining, by a computer, an audio signal including synthetic speech;

extracting, by the computer, metadata from a watermark of the audio signal by applying a set of keys associated with a plurality of text-to-speech (TTS) services to the audio signal, the metadata indicating an origin of the synthetic speech in the audio signal; and

generating, by the computer, based on the metadata as extracted from the watermark, a notification indicating that the audio signal includes the synthetic speech.

2 . The computer-implemented method of claim 1 , further comprising generating a score for each key of the set of keys to determine that the audio signal includes the watermark, wherein the watermark was generated using the key of the set of keys.

3 . The computer-implemented method of claim 2 , further comprising transmitting the key to a TTS service to generate the watermark.

4 . The computer-implemented method of claim 1 , wherein the metadata includes one or more of a service identifier of a TTS service, a model identifier of a TTS model, a user identifier of a user of the TTS service, or a timestamp indicating when the synthetic speech was generated.

5 . The computer-implemented method of claim 1 , further comprising transmitting an alert to a TTS service based on the origin of the synthetic speech in the audio signal.

6 . The computer-implemented method of claim 1 , wherein the notification includes a portion of the metadata as extracted from the watermark.

7 . The computer-implemented method of claim 1 , further comprising:

receiving, by the computer, from the origin of the synthetic speech, the audio signal including the watermark;

determining, by the computer, that a robustness of the watermark exceeds a predetermined threshold; and

transmitting an approval of the watermark to the origin of the synthetic speech.

8 . The computer-implemented method of claim 1 , wherein the watermark includes a consent watermark, and wherein the notification indicates usage consent parameters of the consent watermark.

9 . The computer-implemented method of claim 1 , wherein the watermark includes an authorization watermark, and wherein the notification indicates authorization parameters of the authorization watermark.

10 . The computer-implemented method of claim 1 , further comprising:

obtaining, by the computer, a second audio signal including second synthetic speech;

extracting, by the computer, second metadata from a second watermark of the second audio signal, the second metadata indicating a second origin of the second synthetic speech that is different from the origin of the audio signal including the synthetic speech; and

generating, by the computer, a second notification indicating that the second audio signal includes the second synthetic speech.

11 . A system comprising:

a computing device comprising at least one processor, configured to:

obtain an audio signal including synthetic speech;

extract metadata from a watermark of the audio signal by applying a set of keys associated with a plurality of text-to-speech (TTS) services to the audio signal, the metadata indicating an origin of the synthetic speech in the audio signal; and

generate, based on the metadata as extracted from the watermark, a notification indicating that the audio signal includes the synthetic speech.

12 . The system of claim 11 , wherein the computing device is further configured to generate a score for each key of the set of keys to determine that the audio signal includes the watermark, the key used to generate the watermark.

13 . The system of claim 12 , wherein the computing device is further configured to transmit the key to a TTS service to generate the watermark.

14 . The system of claim 11 , wherein the metadata includes one or more of a service identifier of a TTS service, a model identifier of a TTS model, a user identifier of a user of the TTS service, or a timestamp indicating when the synthetic speech was generated.

15 . The system of claim 11 , wherein the computing device is configured to transmit an alert to a TTS service based on the origin of the synthetic speech in the audio signal.

16 . The system of claim 11 , wherein the notification includes a portion of the metadata as extracted from the watermark.

17 . The system of claim 11 , wherein the computing device is configured to:

receive from the origin of the synthetic speech, the audio signal including the watermark;

determine that a robustness of the watermark exceeds a predetermined threshold; and

transmit an approval of the watermark to the origin of the synthetic speech.

18 . The system of claim 11 , wherein the watermark includes a consent watermark, and wherein the notification indicates one or more usage consent parameters of the consent watermark.

19 . The system of claim 11 , wherein the watermark includes an authorization watermark, and wherein the notification indicates one or more authorization parameters of the authorization watermark.

20 . The system of claim 11 , wherein the computing device is configured to:

obtain a second audio signal including second synthetic speech;

extract second metadata from a second watermark of the second audio signal, the second metadata indicating a second origin of the second synthetic speech that is different from the origin of the audio signal including the synthetic speech; and

generate a second notification indicating that the second audio signal includes the second synthetic speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2024
From: LOONEY, DAVID; GAUBITCH, NIKOLAY; KHOURY, ELIE
To: PINDROP SECURITY, INC.
Reel/Frame 068025/0960 →