IP Library Granted Patent US 12,706,098
Granted Patent B2
US 12,706,098 · App. 18/646,228 · Granted Aug 11, 2026

Active voice liveness detection system

Inventors: Elie Khoury (Atlanta, GA); Ganesh Sivaraman (Atlanta, GA); Tianxiang Chen (Atlanta, GA); Nikolay Gaubitch (Atlanta, GA); David Looney (Atlanta, GA); Amit Gupta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Nicholas Klein (Atlanta, GA); Anthony Stankus (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/26G06F21/32G10L15/02G10L17/00G10L17/02G10L17/04G10L17/08G10L17/18G10L25/51G10L25/60H04M3/2281H04M3/42221H04M3/5175H04M3/5183G10L25/30H04M2201/405
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,706,098
App. No.
18/646,228
Granted
Aug 11, 2026
Kind
B2
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. Embodiments include systems and methods for detecting fraudulent presentation attacks using multiple functional engines that implement various fraud-detection techniques, to produce calibrated scores and/or fused scores. A computer may, for example, evaluate the audio quality of speech signals within audio signals, where speech signals contain the speech portions having speaker utterances.

Claims (61)

1 . A computer-implemented method for detecting machine-based speech in calls, comprising:

obtaining, by a computer, a verification prompt comprising challenge content for display at a user interface of a user device of a speaker;

obtaining, by the computer, an input audio signal comprising a speech signal containing response content as an utterance of the speaker, wherein the response content in the speech signal purportedly matches to the challenge content of the verification prompt;

extracting, by the computer, a text embedding using a first set of features extracted for the text of the challenge content of the verification prompt, and a spoken content embedding using a second set of features extracted using the speech signal of the input audio signal; and

executing, by the computer, a content verification engine to generate a content verification score indicating a probability that the response content matches to the challenge content, the content verification engine having one or more layers of machine-learning architecture trained to determine a distance between the text embedding and a spoken content embedding and output the content verification score according to the distance.

2 . The method according to claim 1 , further comprising:

executing, by the computer, the content verification engine taking the training text embedding and the training response content embedding to generate a predicted content verification score according to a predicted distance between a training text embedding and a training spoken content embedding; and

determining, by the computer, a level of error for the content verification engine based upon an expected content verification score and the predicted content verification score; and

training, by the computer, the content verification engine by updating a set of one or more hyperparameters of the content verification engine according to the level of error.

3 . The method according to claim 2 , further comprising:

obtaining, by the computer, a plurality of training verification prompts and a plurality of training speech signals, wherein each particular training speech signal includes training spoken content that matches to training challenge content of a training verification prompt corresponding to the particular training speech signal.

4 . The method according to claim 2 , further comprising:

executing, by the computer, a text embedding extractor comprising one or more layers of the machine-learning architecture taking the training challenge content of the training verification prompt as input, to extract the training text embedding for the training challenge content; and

training, by the computer, the text embedding extractor by updating a second set of one or more hyperparameters of the text embedding extractor according to one or more levels of error,

wherein the computer executes the text embedding extractor as trained to extract the text embedding for the text of the challenge content of the verification prompt.

5 . The method according to claim 2 , further comprising:

executing, by the computer, a content embedding extractor comprising one or more layers of the machine-learning architecture taking the training spoken content of the training response content of the training speech signal, to extract the training content embedding for the training response content; and

training, by the computer, the content embedding extractor by updating a second set of one or more hyperparameters of the content embedding extractor according to one or more levels of error,

wherein the computer executes the content embedding extractor as trained to extract the spoken content embedding for the text of the spoken response content of the speech signal.

6 . The method according to claim 1 , wherein obtaining the verification prompt comprising the challenge content for display at the user interface of the user device includes:

randomly generating, by the computer, the text of the challenge content of the verification prompt; and

transmitting, by the computer, the verification prompt to the user device.

7 . The method according to claim 1 , wherein obtaining the verification prompt comprising the challenge content for display at the user interface of the user device includes:

selecting, by the computer, from a database the text of the challenge content associated with the speaker associated with the user device for the verification prompt.

8 . The method according to claim 1 , wherein obtaining the input audio signal comprising the speech signal containing the response content as the utterance of a speaker includes:

identifying, by the computer, the utterance occurring in a speech portion of the input audio signal;

generating, by the computer, the speech signal comprising the speech portion of the input audio signal having the utterance, speech portion is filtered away a plurality of non-speech portions of the input audio signal; and

extracting, by the computer, the second set of one or more features for the spoken content embedding from the speech signal of the input audio signal.

9 . The method according to claim 1 , wherein the computer calibrates the content verification score according to one or more quality parameter values associated with the quality parameters.

10 . The method according to claim 1 , further comprising identifying, by the computer, the inbound audio signal as genuine or fraudulent based upon comparing the content verification score against a probability threshold.

11 . A system for detecting machine-based speech in calls, comprising:

a computer comprising at least one processor, configured to:

obtain a verification prompt comprising challenge content for display at a user interface of a user device of a speaker;

obtain an input audio signal comprising a speech signal containing response content as an utterance of the speaker, wherein the response content in the speech signal purportedly matches to the challenge content of the verification prompt;

extract a text embedding using a first set of features extracted for the text of the challenge content of the verification prompt, and a spoken content embedding using a second set of features extracted using the speech signal of the input audio signal; and

execute the content verification engine taking the training text embedding and the training response content embedding to generate a predicted content verification score according to a predicted distance between a training text embedding and a training spoken content embedding.

12 . The system according to claim 11 , wherein the computer is further configured to:

execute the content verification engine taking the training text embedding and the training response content embedding to generate a predicted content verification score according to a predicted distance between a training text embedding and a training spoken content embedding; and

determine a level of error for the content verification engine based upon an expected content verification score and the predicted content verification score; and

train the content verification engine by updating a set of one or more hyperparameters of the content verification engine according to the level of error.

13 . The system according to claim 12 , wherein the computer is further configured to:

obtain a plurality of training verification prompts and a plurality of training speech signals, wherein each particular training speech signal includes training spoken content that matches to training challenge content of a training verification prompt corresponding to the particular training speech signal.

14 . The system according to claim 12 , wherein the computer is further configured to:

execute a text embedding extractor comprising one or more layers of the machine-learning architecture taking the training challenge content of the training verification prompt as input, to extract the training text embedding for the training challenge content; and

train the text embedding extractor by updating a second set of one or more hyperparameters of the text embedding extractor according to one or more levels of error, and

wherein the computer executes the text embedding extractor as trained to extract the text embedding for the text of the challenge content of the verification prompt.

15 . The system according to claim 12 , wherein the computer is further configured to:

execute a content embedding extractor comprising one or more layers of the machine-learning architecture taking the training spoken content of the training response content of the training speech signal, to extract the training content embedding for the training response content; and

train the content embedding extractor by updating a second set of one or more hyperparameters of the content embedding extractor according to one or more levels of error, and

wherein the computer executes the content embedding extractor as trained to extract the spoken content embedding for the text of the spoken response content of the speech signal.

16 . The system according to claim 11 , wherein when obtaining the verification prompt comprising the challenge content for display at the user interface of the user device, the computer is further configured to:

randomly generate the text of the challenge content of the verification prompt; and

transmit the verification prompt to the user device.

17 . The system according to claim 11 , wherein when obtaining the verification prompt comprising the challenge content for display at the user interface of the user device, the computer is further configured to:

select from a database the text of the challenge content associated with the speaker associated with the user device for the verification prompt.

18 . The system according to claim 11 , wherein when obtaining the input audio signal comprising the speech signal containing the response content as the utterance of a speaker the computer is figure configured to:

identify the utterance occurring in a speech portion of the input audio signal;

generate the speech signal comprising the speech portion of the input audio signal having the utterance, speech portion is filtered away a plurality of non-speech portions of the input audio signal; and

extract the second set of one or more features for the spoken content embedding from the speech signal of the input audio signal.

19 . The system according to claim 11 , wherein the computer calibrates the content verification score according to one or more quality parameter values associated with the quality parameters.

20 . The system according to claim 11 , wherein the computer is further configured to identify the inbound audio signal as genuine or fraudulent based upon comparing the content verification score against a probability threshold.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2024
From: SIVARAMAN, GANESH; CHEN, TIANXIANG; GAUBITCH, NIKOLAY; LOONEY, DAVID; KHOURY, ELIE; GUPTA, AMIT; BALASUBRAMANIYAN, VIJAY; KLEIN, NICHOLAS; STANKUS, ANTHONY
To: PINDROP SECURITY, INC.
Reel/Frame 067705/0899 →
Continuity (3)
Provisional Application 63462913 · Apr 28, 2023
Provisional Application 63620068 · Jan 11, 2024
Related Publication 20240363100A1 · Oct 31, 2024
References Cited (141)
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9318114B2 · Zeljkovic · 2016 [cited by examiner]
US 10565498B1 · Zhiyanov · 2020 [cited by applicant]
US 10607148B1 · Niewczas · 2020 [cited by applicant]
US 10657971B1 · Newstadt et al. · 2020 [cited by applicant]
US 10720165B2 · Guo · 2020 [cited by examiner]
US 10853579B2 · Laxman et al. · 2020 [cited by applicant]
US 10878174B1 · Vontobel et al. · 2020 [cited by applicant]
US 11106690B1 · Dhillon et al. · 2021 [cited by applicant]
US 11151385B2 · Iyer et al. · 2021 [cited by applicant]
US 11417339B1 · Wang et al. · 2022 [cited by applicant]
US 11481434B1 · Venti et al. · 2022 [cited by applicant]
US 11532301B1 · Hajebi et al. · 2022 [cited by applicant]
US 11862177B2 · Chen et al. · 2024 [cited by applicant]
US 11943387B1 · Wolinsky et al. · 2024 [cited by applicant]
US 11972760B1 · Chavez et al. · 2024 [cited by applicant]
US 12026241B2 · Lesso · 2024 [cited by applicant]
US 12062105B2 · Liberman et al. · 2024 [cited by applicant]
US 12087319B1 · Looney et al. · 2024 [cited by applicant]
US 12206820B2 · Dropuljic et al. · 2025 [cited by applicant]
US 20030115191A1 · Copperman et al. · 2003 [cited by applicant]
US 20050240779A1 · Aull et al. · 2005 [cited by applicant]
US 20080263661A1 · Bouzida · 2008 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by applicant]
US 20090319274A1 · Gross · 2009 [cited by applicant]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20120215776A1 · Guha et al. · 2012 [cited by applicant]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20140270114A1 · Kolbegger et al. · 2014 [cited by applicant]
US 20150278341A1 · Shen · 2015 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150347902A1 · Butler et al. · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios et al. · 2016 [cited by applicant]
US 20160323398A1 · Guo et al. · 2016 [cited by applicant]
US 20160328547A1 · Gross · 2016 [cited by applicant]
US 20180077101A1 · Desouza Sana et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180261227A1 · Blouet · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190075123A1 · Smith et al. · 2019 [cited by applicant]
US 20190087870A1 · Gardyne et al. · 2019 [cited by applicant]
US 20190114496A1 · Lesso · 2019 [cited by applicant]
US 20190114497A1 · Lesso · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190206423A1 · Winton et al. · 2019 [cited by applicant]
US 20190311722A1 · Caldwell · 2019 [cited by applicant]
US 20190378520A1 · Chiu · 2019 [cited by applicant]
US 20200044852A1 · Streit · 2020 [cited by applicant]
US 20200118544A1 · Lee et al. · 2020 [cited by applicant]
US 20200145816A1 · Morin et al. · 2020 [cited by applicant]
US 20200153535A1 · Jayaweera Kankanamge et al. · 2020 [cited by applicant]
US 20200184979A1 · Keret · 2020 [cited by examiner]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta et al. · 2020 [cited by applicant]
US 20200314246A1 · Hart · 2020 [cited by applicant]
US 20200344251A1 · Jeyakumar et al. · 2020 [cited by applicant]
US 20200366671A1 · Larson et al. · 2020 [cited by applicant]
US 20210110813A1 · Khoury et al. · 2021 [cited by applicant]
US 20210125619A1 · Espejo et al. · 2021 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by applicant]
US 20210168238A1 · Adolphe et al. · 2021 [cited by applicant]
US 20210182584A1 · Ionita · 2021 [cited by applicant]
US 20210193174A1 · Enzinger et al. · 2021 [cited by applicant]
US 20210233541A1 · Chen · 2021 [cited by examiner]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20210312906A1 · Kuo et al. · 2021 [cited by applicant]
US 20210326421A1 · Khoury et al. · 2021 [cited by applicant]
US 20210335354A1 · Park · 2021 [cited by applicant]
US 20210406568A1 · Liberman et al. · 2021 [cited by applicant]
US 20220006899A1 · Phatak et al. · 2022 [cited by applicant]
US 20220027648A1 · Trundle et al. · 2022 [cited by applicant]
US 20220059121A1 · Rao et al. · 2022 [cited by applicant]
US 20220060578A1 · Kee et al. · 2022 [cited by applicant]
US 20220076077A1 · Reddy et al. · 2022 [cited by applicant]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220108701A1 · Gupta et al. · 2022 [cited by applicant]
US 20220124195A1 · Lu et al. · 2022 [cited by applicant]
US 20220156376A1 · Dos Santos Silva et al. · 2022 [cited by applicant]
US 20220215845A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269761A1 · Steelberg et al. · 2022 [cited by applicant]
US 20220269922A1 · Mathews · 2022 [cited by applicant]
US 20220328050A1 · Hennig et al. · 2022 [cited by applicant]
US 20220366901A1 · Rathaur et al. · 2022 [cited by applicant]
US 20220366916A1 · Dos Santos et al. · 2022 [cited by applicant]
US 20230008613A1 · Do et al. · 2023 [cited by applicant]
US 20230032728A1 · Hao · 2023 [cited by applicant]
US 20230082094A1 · Keret et al. · 2023 [cited by applicant]
US 20230107624A1 · Keith · 2023 [cited by applicant]
US 20230134791A1 · Londeree · 2023 [cited by applicant]
US 20230153815A1 · Blouet et al. · 2023 [cited by applicant]
US 20230161853A1 · Nair et al. · 2023 [cited by applicant]
US 20230169370A1 · Elmasry et al. · 2023 [cited by applicant]
US 20230179823A1 · Marten et al. · 2023 [cited by applicant]
US 20230196396A1 · Doumar et al. · 2023 [cited by applicant]
US 20230260521A1 · Slocum et al. · 2023 [cited by applicant]
US 20230262160A1 · Trivedi et al. · 2023 [cited by applicant]
US 20230273978A1 · Suhanic et al. · 2023 [cited by applicant]
US 20230274758A1 · Markhasin et al. · 2023 [cited by applicant]
US 20230308465A1 · Alroobaea et al. · 2023 [cited by applicant]
US 20230386506A1 · Shor et al. · 2023 [cited by applicant]
US 20240040035A1 · Dropuljic et al. · 2024 [cited by applicant]
US 20240040066A1 · Liu et al. · 2024 [cited by applicant]
US 20240073219A1 · Maizels et al. · 2024 [cited by applicant]
US 20240127630A1 · Michaeli et al. · 2024 [cited by applicant]
US 20240127825A1 · Carroll et al. · 2024 [cited by applicant]
US 20240127826A1 · Wolfston et al. · 2024 [cited by applicant]
US 20240203408A1 · Li et al. · 2024 [cited by applicant]
US 20240249712A1 · Yao et al. · 2024 [cited by applicant]
US 20240296698A1 · Matias et al. · 2024 [cited by applicant]
US 20240346850A1 · Kolla et al. · 2024 [cited by applicant]
US 20240363100A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363119A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363123A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363124A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363125A1 · Khoury et al. · 2024 [cited by applicant]
US 20240371510A1 · Aman · 2024 [cited by applicant]
US 20240378274A1 · Tanski et al. · 2024 [cited by applicant]
US 20240406227A1 · Pandit et al. · 2024 [cited by applicant]
US 20240420725A1 · Caven et al. · 2024 [cited by applicant]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by applicant]
US 20250078841A1 · Wu et al. · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar et al. · 2025 [cited by applicant]
US 20250273232A1 · Seo et al. · 2025 [cited by applicant]
EP 3156978A1 · 2017 [cited by applicant]
FR 3105479A1 · 2021 [cited by applicant]
GB 2612032A · 2023 [cited by applicant]
WO WO2008047339A2 · 2008 [cited by applicant]
WO WO2021154600A1 · 2021 [cited by applicant]
WO WO2022082036A1 · 2022 [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US 24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
Mittal et al. “AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]
Shiota et al., “Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification” published in Interspeech 2015 16th Annual Conference of the International Speech Communic… [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/388,428 dated Apr. 10, 2026. [cited by applicant]
Al-Rahman et al., “Using Deep Learning Neural Networks to Recognize and Authenticate the Identity of the Speaker”, 2024 21st International Multi-Conference on Systems, Signals & Devices (SSD). (8 pages). [cited by applicant]
Fazel et al., “ An Overview of Statistical Pattern Recognition Techniques for Speaker Verification”, IEEE Circuits and Systems Magazine, Second Quarter 2011. (20 pages). [cited by applicant]
Shayamunda et al., “Biometric Authentication System for Industrial Applications using Speaker Recognition”, Department of Electrical, Electronic and Computer Engineering, University of Pretoria Tshwane, South Africa, 20… [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/646,431, dated Jul. 15, 2026. [cited by applicant]