IP Library Granted Patent US 12,592,239
Granted Patent B2
US 12,592,239 · App. 18/646,310 · Granted Mar 31, 2026

Active voice liveness detection system

Inventors: Elie Khoury (Atlanta, GA); Ganesh Sivaraman (Atlanta, GA); Tianxiang Chen (Atlanta, GA); Nikolay Gaubitch (Atlanta, GA); David Looney (Atlanta, GA); Amit Gupta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Nicholas Klein (Atlanta, GA); Anthony Stankus (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/26G06F21/32G10L15/02G10L17/00G10L17/02G10L17/04G10L17/08G10L17/18G10L25/51G10L25/60H04M3/2281H04M3/42221H04M3/5175H04M3/5183G10L25/30H04M2201/405
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,239
App. No.
18/646,310
Granted
Mar 31, 2026
Kind
B2
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.

Claims (50)

1 . A computer-implemented method for detecting machine-based speech in calls, comprising:

obtaining, by a computer, an inbound audio signal comprising a speech signal containing response content as an utterance of the speaker, wherein the response content in the speech signal purportedly matches to challenge content of a verification prompt;

extracting, by the computer, a text embedding using a first set of features extracted for text of the challenge content, a spoken content embedding using a second set of features extracted for the speech signal, and a fakeprint using a third set one or more features extracted for one or more fraud artifacts of the speech signal;

generating, by the computer, a content verification score based upon a distance between the text embedding and the spoken content embedding;

executing, by the computer, a passive liveness detector to generate a passive liveness score for the inbound audio signal, the passive liveness detector having a set of layers of a machine-learning architecture trained to classify and score the input audio signal based upon the fakeprint extracted for the fraud artifacts of the inbound audio signal;

generating, by the computer, a fused liveness score based upon the content verification score and the passive liveness score; and

identifying, by the computer, the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold.

2 . The method according to claim 1 , further comprising:

extracting, by the computer, an inbound voiceprint for the speech signal using a fourth set of one or more features extracted for one or more acoustic features of the speech signal of the inbound audio signal; and

generating, by the computer, a speaker verification score for the speech signal indicating a speaker recognition likelihood that the speaker is an enrolled user based upon a second distance between the inbound voiceprint and an enrolled voiceprint.

3 . The method according to claim 2 , wherein the computer generates the fused liveness score further using the speaker verification score.

4 . The method according to claim 1 , further comprising:

generating, by the computer, one or more acoustic parameters corresponding to one or more types of degradation in the speech signal of the inbound audio signal; and

generating, by the computer, a speech quality score for the speech signal based upon the one or more acoustic parameters.

5 . The method according to claim 4 , wherein generating the content verification score includes:

calibrating, by the computer, the content verification score based upon the speech quality score.

6 . The method according to claim 4 , further comprising:

determining, by the computer, that the speech quality score for the speech signal fails a speech quality threshold; and

transmitting, by the computer, to the user device a request for an improved speech signal for the caller.

7 . The method according to claim 1 , further comprising:

extracting, by the computer, an inbound audioprint using one or more features extracted from the audio signal;

generating, by the computer, an audio replay score for the inbound audio signal indicating an audio recording recognition likelihood that the inbound audio signal matches a prior audio signal based upon a distance between the inbound audioprint and a stored audioprint for the prior audio signal.

8 . The method according to claim 7 , further comprising identifying, by the computer, the inbound audio signal as fraudulent, in response to determining that the audio replay score satisfies a replay detection threshold value.

9 . The method according to claim 7 , further comprising storing, by the computer, the inbound audioprint into a database as a new stored audioprint.

10 . The method according to claim 1 , further comprising generating, by the computer, a verification prompt including the challenge content for display at a user interface of the user device associated with the caller.

11 . A system for detecting machine-based speech in calls, comprising:

a computer having at least one processor, configured to:

obtain an inbound audio signal comprising a speech signal containing response content as an utterance of the speaker, wherein the response content in the speech signal purportedly matches to challenge content of a verification prompt;

extract a text embedding using a first set of features extracted for text of the challenge content, a spoken content embedding using a second set of features extracted for the speech signal, and a fakeprint using a third set one or more features extracted for one or more fraud artifacts of the speech signal;

generate a content verification score based upon a distance between the text embedding and the spoken content embedding;

execute a passive liveness detector having a set of layers of a machine-learning architecture to generate a passive liveness score for the inbound audio signal, the passive liveness detector trained to classify and score the input audio signal based upon the fakeprint extracted for the fraud artifacts of the inbound audio signal;

generate a fused liveness score based upon the content verification score and the passive liveness score; and

identify the inbound audio signal as genuine or fraudulent based upon comparing the fused liveness score against an overall risk threshold.

12 . The system according to claim 11 , wherein the computer is further configured to:

extract an inbound voiceprint for the speech signal using a fourth set of one or more features extracted for one or more acoustic features of the speech signal of the inbound audio signal; and

generate a speaker verification score for the speech signal indicating a speaker recognition likelihood that the speaker is an enrolled user based upon a second distance between the inbound voiceprint and an enrolled voiceprint.

13 . The system according to claim 12 , wherein the computer generates the fused liveness score further using the speaker verification score.

14 . The system according to claim 11 , wherein the computer is further configured to:

generate one or more acoustic parameters corresponding to one or more types of degradation in the speech signal of the inbound audio signal; and

generate a speech quality score for the speech signal based upon the one or more acoustic parameters.

15 . The system according to claim 14 , wherein when generating the content verification score the computer is further configured to calibrate the content verification score based upon the speech quality score.

16 . The system according to claim 14 , wherein the computer is further configured to:

determine that the speech quality score for the speech signal fails a speech quality threshold; and

transmit to the user device a request for an improved speech signal for the caller.

17 . The system according to claim 11 , wherein the computer is further configured to:

extract an inbound audioprint using one or more features extracted from the audio signal;

generate an audio replay score for the inbound audio signal indicating an audio recording recognition likelihood that the inbound audio signal matches a prior audio signal based upon a distance between the inbound audioprint and a stored audioprint for the prior audio signal.

18 . The system according to claim 17 , wherein the computer is further configured to identify the inbound audio signal as fraudulent, in response to determining that the audio replay score satisfies a replay detection threshold value.

19 . The system according to claim 17 , wherein the computer is further configured to store the inbound audioprint into a database as a new stored audioprint.

20 . The system according to claim 11 , wherein the computer is further configured to generate a verification prompt including the challenge content for display at a user interface of the user device associated with the caller.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2024
From: SIVARAMAN, GANESH; CHEN, TIANXIANG; GAUBITCH, NIKOLAY; LOONEY, DAVID; KHOURY, ELIE; GUPTA, AMIT; BALASUBRAMANIYAN, VIJAY; KLEIN, NICHOLAS; STANKUS, ANTHONY
To: PINDROP SECURITY, INC.
Reel/Frame 067707/0604 →
Continuity (3)
Provisional Application 63620068 · Jan 11, 2024
Provisional Application 63462913 · Apr 28, 2023
Related Publication 20240363123A1 · Oct 31, 2024
References Cited (107)
US 9060057B1 · Danis · 2015 [cited by examiner]
US 9318114B2 · Zeljkovic et al. · 2016 [cited by applicant]
US 10565498B1 · Zhiyanov · 2020 [cited by applicant]
US 10607148B1 · Niewczas · 2020 [cited by examiner]
US 10657971B1 · Newstadt · 2020 [cited by examiner]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 10853579B2 · Laxman et al. · 2020 [cited by applicant]
US 10878174B1 · Vontobel et al. · 2020 [cited by applicant]
US 11106690B1 · Dhillon et al. · 2021 [cited by applicant]
US 11151385B2 · Iyer · 2021 [cited by examiner]
US 11417339B1 · Wang · 2022 [cited by examiner]
US 11481434B1 · Venti et al. · 2022 [cited by applicant]
US 11532301B1 · Hajebi et al. · 2022 [cited by applicant]
US 11943387B1 · Wolinsky et al. · 2024 [cited by applicant]
US 11972760B1 · Chavez et al. · 2024 [cited by applicant]
US 12026241B2 · Lesso · 2024 [cited by applicant]
US 12062105B2 · Liberman · 2024 [cited by examiner]
US 12087319B1 · Looney et al. · 2024 [cited by applicant]
US 20030115191A1 · Copperman et al. · 2003 [cited by applicant]
US 20050240779A1 · Aull et al. · 2005 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by examiner]
US 20090319274A1 · Gross · 2009 [cited by examiner]
US 20120215776A1 · Guha et al. · 2012 [cited by applicant]
US 20130218566A1 · Qian · 2013 [cited by examiner]
US 20130262096A1 · Wilhelms-Tricarico · 2013 [cited by examiner]
US 20140270114A1 · Kolbegger · 2014 [cited by examiner]
US 20150278341A1 · Shen · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios · 2016 [cited by examiner]
US 20160323398A1 · Guo et al. · 2016 [cited by applicant]
US 20160328547A1 · Gross · 2016 [cited by examiner]
US 20180077101A1 · Desouza Sana et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury · 2018 [cited by examiner]
US 20180261227A1 · Blouet · 2018 [cited by examiner]
US 20190087870A1 · Gardyne · 2019 [cited by examiner]
US 20190114496A1 · Lesso · 2019 [cited by examiner]
US 20190114497A1 · Lesso · 2019 [cited by applicant]
US 20190174000A1 · Bharrat · 2019 [cited by examiner]
US 20190206423A1 · Winton · 2019 [cited by examiner]
US 20190311722A1 · Caldwell · 2019 [cited by examiner]
US 20190378520A1 · Chiu · 2019 [cited by examiner]
US 20200044852A1 · Streit · 2020 [cited by examiner]
US 20200118544A1 · Lee et al. · 2020 [cited by applicant]
US 20200184979A1 · Keret et al. · 2020 [cited by applicant]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta · 2020 [cited by examiner]
US 20200314246A1 · Hart · 2020 [cited by examiner]
US 20210125619A1 · Espejo et al. · 2021 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by examiner]
US 20210168238A1 · Adolphe · 2021 [cited by examiner]
US 20210182584A1 · Ionita · 2021 [cited by applicant]
US 20210193174A1 · Enzinger · 2021 [cited by examiner]
US 20210233541A1 · Chen · 2021 [cited by examiner]
US 20210280171A1 · Phatak · 2021 [cited by examiner]
US 20210312906A1 · Kuo · 2021 [cited by examiner]
US 20210335354A1 · Park · 2021 [cited by examiner]
US 20210406568A1 · Liberman et al. · 2021 [cited by applicant]
US 20220027648A1 · Trundle · 2022 [cited by examiner]
US 20220060578A1 · Kee · 2022 [cited by examiner]
US 20220076077A1 · Reddy · 2022 [cited by examiner]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220108701A1 · Gupta · 2022 [cited by examiner]
US 20220124195A1 · Lu · 2022 [cited by examiner]
US 20220215845A1 · Sharifi · 2022 [cited by examiner]
US 20220269922A1 · Mathews · 2022 [cited by examiner]
US 20220328050A1 · Hennig · 2022 [cited by examiner]
US 20220366901A1 · Rathaur et al. · 2022 [cited by applicant]
US 20220366916A1 · dos Santos · 2022 [cited by examiner]
US 20230008613A1 · Do · 2023 [cited by examiner]
US 20230032728A1 · Hao · 2023 [cited by examiner]
US 20230082094A1 · Keret · 2023 [cited by examiner]
US 20230107624A1 · Keith · 2023 [cited by applicant]
US 20230134791A1 · Londeree · 2023 [cited by applicant]
US 20230153815A1 · Blouet · 2023 [cited by examiner]
US 20230161853A1 · Nair · 2023 [cited by examiner]
US 20230179823A1 · Marten · 2023 [cited by examiner]
US 20230196396A1 · Doumar · 2023 [cited by examiner]
US 20230260521A1 · Slocum et al. · 2023 [cited by applicant]
US 20230262160A1 · Trivedi · 2023 [cited by examiner]
US 20230274758A1 · Markhasin · 2023 [cited by examiner]
US 20230386506A1 · Shor · 2023 [cited by examiner]
US 20240040035A1 · Dropuljic et al. · 2024 [cited by applicant]
US 20240040066A1 · Liu · 2024 [cited by examiner]
US 20240073219A1 · Maizels · 2024 [cited by examiner]
US 20240127630A1 · Michaeli · 2024 [cited by examiner]
US 20240127826A1 · Wolfston et al. · 2024 [cited by applicant]
US 20240203408A1 · Li · 2024 [cited by examiner]
US 20240249712A1 · Yao · 2024 [cited by examiner]
US 20240296698A1 · Matias · 2024 [cited by examiner]
US 20240346850A1 · Kolla · 2024 [cited by examiner]
US 20240371510A1 · Aman · 2024 [cited by examiner]
US 20240378274A1 · Tanski et al. · 2024 [cited by applicant]
US 20240406227A1 · Pandit · 2024 [cited by examiner]
US 20240420725A1 · Caven, III · 2024 [cited by examiner]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by applicant]
US 20250078841A1 · Wu et al. · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar et al. · 2025 [cited by applicant]
US 20250273232A1 · Seo · 2025 [cited by examiner]
EP 3156978A1 · 2017 [cited by applicant]
FR 3105479A1 · 2021 [cited by applicant]
GB 2612032A · 2023 [cited by applicant]
WO WO2008047339A2 · 2008 [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US 24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
Shiota et al., “Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification” published in Interspeech 2015 16th Annual Conference of the International Speech Communic… [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
Mittal et al. “AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]