IP Library Patent Application 18646493
Patent Application
App. No. 18/646,493

ACTIVE VOICE LIVENESS DETECTION SYSTEM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/646,493
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.

Claims (41)

1 . A computer-implemented method for generating liveness scores for detecting fraud occurring in calls, comprising:

obtaining, by a computer, an input audio signal including one or more speech signals representing one or more utterances of a speaker;

extracting, by the computer, a fakeprint for the input audio signal using one or more fraud artifact features extracted from the input audio signal;

determining, by the computer, a magnitude value for the fakeprint based upon a vector length of the fakeprint;

executing, by the computer, a passive liveness detector having one or more layers of a machine-learning architecture to generate a liveness score for the input audio signal, the passive liveness detector trained to determine the liveness score taking the fakeprint as an input and calibrate the liveness score using the magnitude value.

2 . The method according to claim 1 , further comprising:

determining, by the computer, a net speech value for the speaker indicating an aggregate amount of each utterance obtained for the speaker,

wherein the passive liveness detector is trained to calibrate the liveness score further using the net speech value.

3 . The method according to claim 1 , further comprising:

determining, by the computer, a plurality of audio quality parameters for the input audio signal, including the magnitude value and at least one of a net speech value, SNR, DRR, T60, or C50,

wherein the passive liveness detector is trained to calibrate the liveness score further using the plurality of audio quality parameters.

4 . The method according to claim 3 , further comprising:

executing, by the computer, a content verifier of the machine-learning architecture to generate a content verification score for the input audio signal, the content verifier trained to determine the content verification score taking challenge content and response content and calibrate the content verification score using one or more audio quality of parameters.

5 . The method according to claim 3 , further comprising:

executing, by the computer, a speaker verifier of the machine-learning architecture to generate a speaker verification score for the input audio signal, the speaker verifier trained to determine the speaker verification score taking a voiceprint for the inbound audio signal and calibrate the speaker verification score using one or more audio quality of parameters.

6 . The method according to claim 3 , further comprising:

generating, by the computer, a fused liveness score for the input audio signal based upon the liveness score and at least one of a content verification score for the input audio signal or a speaker verification score for the input audio signal.

7 . The method according to claim 1 , wherein the input audio signal includes an inbound audio signal received from a user device that originated an inbound call during a deployment phase,

8 . The method according to claim 1 , wherein the input audio signal includes a training audio signal received from a database during a training phase,

9 . The method according to claim 1 , wherein the input audio signal includes an enrollment audio signal received from a user device during an enrollment phase.

10 . The method according to claim 1 , further comprising parsing, by the computer, the input audio signal into one or more segments, wherein the computer determines the liveness score for each successive segment of the input audio signal.

11 . A system for generating liveness scores for detecting fraud occurring in calls, comprising:

computer comprising a processor configured to:

obtain an input audio signal including one or more speech signals representing one or more utterances of a speaker;

extract a fakeprint for the input audio signal using one or more fraud artifact features extracted from the input audio signal;

determine a magnitude value for the fakeprint based upon a vector length of the fakeprint; and

execute a passive liveness detector to generate a liveness score for the input audio signal, the passive liveness detector having one or more layers of a machine-learning architecture trained to determine the liveness score taking the fakeprint as an input and calibrate the liveness score using the magnitude value.

12 . The system according to claim 1 , wherein the computer is further configured to:

determine a net speech value for the speaker indicating an aggregate amount of each utterance obtained for the speaker,

wherein the passive liveness detector is trained to calibrate the liveness score further using the net speech value.

13 . The system according to claim 1 , wherein the computer is further configured to:

determine a plurality of audio quality parameters for the input audio signal, including the magnitude value and at least one of a net speech value, SNR, DRR, T60, or C50; and

wherein the passive liveness detector is trained to calibrate the liveness score further using the plurality of audio quality parameters.

14 . The system according to claim 3 , wherein the computer is further configured to execute a content verifier of the machine-learning architecture to generate a content verification score for the input audio signal, the content verifier trained to determine the content verification score taking challenge content and response content and calibrate the content verification score using one or more audio quality of parameters.

15 . The system according to claim 3 , wherein the computer is further configured to execute a speaker verifier of the machine-learning architecture to generate a content verification score for the input audio signal, the speaker verifier trained to determine the speaker verification score taking a voiceprint for the inbound audio signal and calibrate the speaker verification score using one or more audio quality of parameters.

16 . The system according to claim 13 , wherein the computer is further configured to generate a fused liveness score for the input audio signal based upon the liveness score and at least one of a content verification score for the input audio signal or a speaker verification score for the input audio signal.

17 . The system according to claim 10 , wherein the input audio signal includes an inbound audio signal received from a user device that originated an inbound call during a deployment phase,

18 . The system according to claim 10 , wherein the input audio signal includes a training audio signal received from a database during a training phase,

19 . The system according to claim 10 , wherein the input audio signal includes an enrollment audio signal received from a user device during an enrollment phase.

20 . The system according to claim 11 , wherein the computer is further configured to:

parse the input audio signal into one or more segments, wherein the computer determines the liveness score for each successive segment of the input audio signal.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2024
From: SIVARAMAN, GANESH; CHEN, TIANXIANG; GAUBITCH, NIKOLAY; LOONEY, DAVID; KHOURY, ELIE; GUPTA, AMIT; BALASUBRAMANIYAN, VIJAY; KLEIN, NICHOLAS; STANKUS, ANTHONY
To: PINDROP SECURITY, INC.
Reel/Frame 067707/0604 →