IP Library Patent Application 18646431
Patent Application
App. No. 18/646,431

ACTIVE VOICE LIVENESS DETECTION SYSTEM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/646,431
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.

Claims (49)

1 . A computer-implemented method for detecting fraudulent speech in media data, comprising:

receiving, by a computer, media data including a speech signal in an audio signal of the media data;

determining, by the computer, a plurality of segments of the media data according to a preconfigured segmenting boundary;

for each segment of the media data,

extracting, by the computer, a segment fakeprint for the segment using a plurality of segment features extracted using a speech portion of the speech signal occurring in the segment;

generating, by the computer, a segment liveness score for the segment based upon a distance between the segment fakeprint and a classification threshold value; and

identifying, by the computer, the portion of the speech signal in the segment as genuine or fraudulent based upon comparing the segment liveness score for the segment against a fraud detection threshold.

2 . The method according to claim 1 , further comprising:

receiving, by a computer, one or more enrollment media data inputs, enrollment media data input including an enrollment speech signal in an enrollment audio signal;

extracting, by the computer, an enrolled fakeprint using a plurality of features extracted using each enrollment speech signal of each enrollment media data input;

determining, by the computer, the distance between the segment fakeprint and the enrolled fakeprint as the classification threshold value.

3 . The method according to claim 1 , further comprising determining, by the computer, the speech portion of the speech signal in the inbound audio signal occurring in the segment, wherein the computer ignores a non-speech portion of the input audio signal in the segment.

4 . The method according to claim 3 , further comprising executing, by the computer, at least one of voice activity detection (VAD), an automatic speech recognition (ASR), or speaker diarization, to determine the speech portion of the speech signal in the inbound audio signal in the segment.

5 . The method according to claim 1 , wherein the segmenting boundary is based upon a uniform time interval.

6 . The method according to claim 1 , wherein the segmenting boundary is based on detecting a triggering condition associated with the speech signal.

7 . The method according to claim 1 , further comprising in response to identifying the speech portion of the speech signal in the segment as fraudulent, generating, by the computer, an alert notification indicating a timestamp of the segment in the media data.

8 . A system for detecting fraudulent speech in media data comprising:

a computer comprising at least one processor configured to:

receive media data including a speech signal in an audio signal of the media data;

determine a plurality of segments of the media data according to a preconfigured segmenting boundary;

for each segment of the media data,

extract a segment fakeprint for the segment using a plurality of segment features extracted using a speech portion of the speech signal occurring in the segment;

generate a segment liveness score for the segment based upon a distance between the segment fakeprint and a classification threshold value; and

identify the portion of the speech signal in the segment as genuine or fraudulent based upon comparing the segment liveness score for the segment against a fraud detection threshold.

9 . The system according to claim 8 , wherein the computer is further configured to:

receive one or more enrollment media data inputs, enrollment media data input including an enrollment speech signal in an enrollment audio signal;

extract an enrolled fakeprint using a plurality of features extracted using each enrollment speech signal of each enrollment media data input;

determine the distance between the segment fakeprint and the enrolled fakeprint as the classification threshold value.

10 . The system according to claim 8 , wherein the computer is further configured to determine the speech portion of the speech signal in the inbound audio signal occurring in the segment, wherein the computer ignores a non-speech portion of the input audio signal in the segment.

11 . The system according to claim 10 , wherein the computer is further configured to execute at least one of voice activity detection (VAD), an automatic speech recognition (ASR), or speaker diarization, to determine the speech portion of the speech signal in the inbound audio signal in the segment.

12 . The system according to claim 8 , wherein the segmenting boundary is based upon at least one of a uniform time interval or in response to detecting a triggering condition.

13 . The system according to claim 8 , wherein the segmenting boundary is based on detecting a triggering condition associated with the speech signal.

14 . The system according to claim 8 , wherein the computer is further configured to, in response to identifying the speech portion of the speech signal in the segment as fraudulent, generate an alert notification indicating a timestamp of the segment in the media data.

15 . A non-transitory computer-readable media configured to store machine-executable instructions that when executed by one or more processors cause the processors to:

receive media data including a speech signal in an audio signal of the media data;

determine a plurality of segments of the media data according to a preconfigured segmenting boundary;

for each segment of the media data,

extract a segment fakeprint for the segment using a plurality of segment features extracted using a speech portion of the speech signal occurring in the segment;

generate a segment liveness score for the segment based upon a distance between the segment fakeprint and a classification threshold value; and

identify the portion of the speech signal in the segment as genuine or fraudulent based upon comparing the segment liveness score for the segment against a fraud detection threshold.

16 . The non-transitory medium of claim 13 , wherein the instructions cause the one or more processors:

receive one or more enrollment media data inputs, enrollment media data input including an enrollment speech signal in an enrollment audio signal;

extract an enrolled fakeprint using a plurality of features extracted using each enrollment speech signal of each enrollment media data input;

determine the distance between the segment fakeprint and the enrolled fakeprint as the classification threshold value.

17 . The non-transitory medium of claim 16 , wherein the instructions cause the one or more processors to:

determine the speech portion of the speech signal in the inbound audio signal occurring in the segment, wherein the computer ignores a non-speech portion of the input audio signal in the segment.

18 . The non-transitory medium of claim 16 , wherein the instructions cause the one or more processors to execute at least one of voice activity detection (VAD), an automatic speech recognition (ASR), or speaker diarization, to determine the speech portion of the speech signal in the inbound audio signal in the segment.

19 . The non-transitory medium of claim 16 , wherein the segmenting boundary is based upon at least one of a uniform time interval or in response to detecting a triggering condition.

20 . The non-transitory medium of claim 16 , wherein the instructions cause the one or more processors to, in response to identifying the speech portion of the speech signal in the segment as fraudulent, generate an alert notification indicating a timestamp of the segment in the media data.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2024
From: SIVARAMAN, GANESH; CHEN, TIANXIANG; GAUBITCH, NIKOLAY; LOONEY, DAVID; KHOURY, ELIE; GUPTA, AMIT; BALASUBRAMANIYAN, VIJAY; KLEIN, NICHOLAS; STANKUS, ANTHONY
To: PINDROP SECURITY, INC.
Reel/Frame 067707/0604 →