IP Library Patent Application 18646375
Patent Application
App. No. 18/646,375

ACTIVE VOICE LIVENESS DETECTION SYSTEM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/646,375
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.

Claims (59)

1 . A computer-implemented method for detecting fraud in calls by repeated recordings, the method comprising:

receiving, by a computer, an inbound audio signal from a user device associated with a caller containing a speech signal for one or more utterances of the caller;

extracting, by the computer, an inbound audioprint for the inbound audio signal using the one or more features extracted from the speech signal of the inbound audio signal;

generating, by the computer, an audio replay score for the inbound audio signal indicating an audio recording recognition likelihood that the inbound audio signal matches a prior audio signal based upon a distance between the inbound audioprint and a prior audioprint for the prior audio signal; and

identifying, by the computer, the inbound audio signal as a replayed recording or unrecognized recording based upon comparing the audio replay score against a replay detection threshold.

2 . The method according to claim 1 , further comprising:

obtaining, by the computer, a plurality of prior audio signals, each prior audio signal comprises a corresponding prior speech signal for one or more prior utterances;

extracting, by the computer, a plurality of audioprints corresponding to the plurality of prior audio signals using the one or more features extracted for the prior speech signal;

storing, by the computer, each prior audioprint into a database.

3 . The method according to claim 2 , further comprising:

executing, by the computer, one or more data augmentation operations using the one or more features extracted from the prior audio signal to generate one or more simulated prior audio signals; and

extracting, by the computer, the one or more features from each simulated prior audio signal of the one or more simulated prior audio signals,

wherein the computer extracts the prior audioprint using the one or more features extracted from the prior audio signal and from the one or more simulated prior audio signals.

4 . The method according to claim 1 , further comprising, in response to identifying the inbound audio signal as an unrecognized recording, storing, by the computer, the inbound audioprint into a database as a next prior audioprint.

5 . The method according to claim 1 , further comprising, in response to identifying the inbound audio signal as a replayed recording, generating, by the computer, an alert notification indicating that the computer detected the replayed recording.

6 . The method according to claim 1 , further comprising:

generating, by the computer, a message indicating that the computer has identified the inbound audio signal as a replayed recording or unrecognized recording; and

transmitting, by the computer, a notification containing the message for presentation at a user interface associated with an agent-user.

7 . The method according to claim 1 , wherein the inbound audio signal is an enrollment audio signal received by the computer during an enrollment of the caller, and wherein the computer halts the enrollment for a preconfigured period of time in response to identifying that the inbound audio signal as the enrollment audio signal is a replayed recording.

8 . The method according to claim 1 , further comprising:

detecting, by the computer, the one or more utterances occurring in one or more speech portions of the input audio signal; and

generating, by the computer, the speech signal comprising the one or more speech portions of the inbound audio signal, wherein the one or more speech portions are filtered away from a plurality of non-speech portions of the input audio signal.

9 . A system for detecting fraud in calls by repeated recordings, the system comprising:

a computer comprising at least one processor, configured to:

receive an inbound audio signal from a user device associated with a caller containing a speech signal for one or more utterances of the caller;

extract an inbound audioprint for the inbound audio signal using the one or more features extracted from the speech signal of the inbound audio signal;

generate an audio replay score for the inbound audio signal indicating an audio recording recognition likelihood that the inbound audio signal matches a prior audio signal based upon a distance between the inbound audioprint and a prior audioprint for the prior audio signal; and

identify the inbound audio signal as a replayed recording or unrecognized recording based upon comparing the audio replay score against a replay detection threshold.

10 . The system according to claim 9 , wherein the computer is further configured to:

obtain a plurality of prior audio signals, each prior audio signal comprises a corresponding prior speech signal for one or more prior utterances;

extract a plurality of audioprints corresponding to the plurality of prior audio signals using the one or more features extracted for the prior speech signal;

store each prior audioprint into a database.

11 . The system according to claim 10 , wherein the computer is further configured to:

execute one or more data augmentation operations using the one or more features extracted from the prior audio signal to generate one or more simulated prior audio signals; and

extract the one or more features from each simulated prior audio signal of the one or more simulated prior audio signals,

wherein the computer extracts the prior audioprint using the one or more features extracted from the prior audio signal and from the one or more simulated prior audio signals.

12 . The system according to claim 9 , wherein the computer is further configured to, in response to identifying the inbound audio signal as an unrecognized recording, storing, by the computer, the inbound audioprint into a database as a next prior audioprint.

13 . The system according to claim 9 , wherein the computer is further configured to, in response to identifying the inbound audio signal as a replayed recording, generating, by the computer, an alert notification indicating that the computer detected the replayed recording.

14 . The system according to claim 9 , wherein the computer is further configured to:

generate a message indicating that the computer has identified the inbound audio signal as a replayed recording or unrecognized recording; and

transmit a notification containing the message for presentation at a user interface associated with an agent-user.

15 . The system according to claim 9 , wherein the inbound audio signal is an enrollment audio signal received by the computer during an enrollment of the caller, and wherein the computer halts the enrollment for a preconfigured period of time in response to identifying that the inbound audio signal as the enrollment audio signal is a replayed recording.

16 . The system according to claim 9 , wherein the computer is further configured to:

detect the one or more utterances occurring in one or more speech portions of the input audio signal;

generate the speech signal comprising the one or more speech portions of the inbound audio signal, wherein the one or more speech portions are filtered away from a plurality of non-speech portions of the input audio signal.

17 . A non-transitory computer-readable media configured to store machine-executable instructions that when executed by one or more processors cause the processors to:

receive an inbound audio signal from a user device associated with a caller containing a speech signal for one or more utterances of the caller;

extract an inbound audioprint for the inbound audio signal using the one or more features extracted from the speech signal of the inbound audio signal;

generate an audio replay score for the inbound audio signal indicating an audio recording recognition likelihood that the inbound audio signal matches a prior audio signal based upon a distance between the inbound audioprint and a prior audioprint for the prior audio signal; and

identify the inbound audio signal as a replayed recording or unrecognized recording based upon comparing the audio replay score against a replay detection threshold.

18 . The non-transitory media of claim 17 , wherein the instructions further cause the one or more processors to:

obtain a plurality of prior audio signals, each prior audio signal comprises a corresponding prior speech signal for one or more prior utterances;

extract a plurality of audioprints corresponding to the plurality of prior audio signals using the one or more features extracted for the prior speech signal; and

store each prior audioprint into a database.

19 . The non-transitory media of claim 18 , wherein the instructions further cause the one or more processors to:

execute one or more data augmentation operations using the one or more features extracted from the prior audio signal to generate one or more simulated prior audio signals; and

extract the one or more features from each simulated prior audio signal of the one or more simulated prior audio signals,

wherein the computer extracts the prior audioprint using the one or more features extracted from the prior audio signal and from the one or more simulated prior audio signals.

20 . The non-transitory media of claim 17 , wherein the instructions further cause the one or more processors to, in response to identifying the inbound audio signal as an unrecognized recording, store the inbound audioprint into a database as a next prior audioprint.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2024
From: SIVARAMAN, GANESH; CHEN, TIANXIANG; GAUBITCH, NIKOLAY; LOONEY, DAVID; KHOURY, ELIE; GUPTA, AMIT; BALASUBRAMANIYAN, VIJAY; KLEIN, NICHOLAS; STANKUS, ANTHONY
To: PINDROP SECURITY, INC.
Reel/Frame 067707/0604 →