IP Library Granted Patent US 12,711,951
Granted Patent B2
US 12,711,951 · App. 18/388,457 · Granted Aug 18, 2026

Deepfake detection

Inventors: Umair Altaf (Atlanta, GA); Sai Pradeep Peri (Atlanta, GA); Lakshay Phatela (Atlanta, GA); Payas Gupta (Atlanta, GA); Yitao Sun (Atlanta, GA); Svetlana Afanaseva (Atlanta, GA); Kailash Patil (Atlanta, GA); Elie Khoury (Atlanta, GA); Bradley Magnetta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Tianxiang Chen (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L15/08G06N20/00G10L15/02G10L15/16G10L15/26G10L17/06G10L17/18G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,711,951
App. No.
18/388,457
Granted
Aug 18, 2026
Kind
B2
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.

Claims (33)

1 . A computer-implemented method for authenticating users based on speech of audio signals, comprising:

receiving, by a computer via a first communication channel for a resource, a user authentication request for accessing the resource including authentication credentials for a first authentication;

obtaining, by the computer via a second communication channel, responsive to successfully authenticating a user according to the first authentication using the authentication credentials, an audio speech signal of the user for a second authentication;

extracting, by the computer, a plurality of features from the audio speech signal, including one or more audio background features;

generating, by the computer, a plurality of initial liveness scores using the plurality of features by applying a plurality of machine-learning models of a machine-learning architecture to the plurality of features extracted from the audio speech signal, wherein each machine-learning model of the machine-learning architecture is trained to generate a corresponding initial liveness score of the plurality of initial liveness scores using a corresponding subset of features of the plurality of features extracted from the audio speech signal, the initial liveness scores including a first initial liveness score indicating an audio background change generated using at least one subset of features having the one or more audio background features from the plurality of features;

generating, by the computer, based upon the plurality of initial liveness scores generated using the plurality of features extracted from the audio speech signal via the second communication channel, a liveness score indicating a likelihood that the user is a human speaker; and

executing, by the computer, the second authentication to determine whether to permit the user access to the resource of the first communication channel based upon comparing the liveness score against a threshold.

2 . The method of claim 1 , wherein obtaining the audio speech signal includes presenting, by the computer, a prompt to direct the user to provide the audio speech signal, responsive to authenticating the user in the first authentication using the authentication credentials.

3 . The method of claim 1 , wherein generating the liveness score includes applying, by the computer, the machine-learning architecture to the plurality of features to generate the liveness score,

wherein the machine-learning architecture is trained on a plurality of examples, each example identifying a second plurality of features and a label indicating one of human speech or machine-generated speech.

4 . The method of claim 1 , wherein the plurality of initial liveness scores further comprises at least one of: a second score identifying passive liveness of the speech signal of the speaker, or a third score identifying repetition of speech within the speech signal.

5 . The method of claim 1 wherein extracting the plurality of features includes applying, by the computer, the machine-learning architecture to generate the plurality of features including a set of embeddings representing spoofing artifacts in the audio speech signal.

6 . The method of claim 5 , wherein generating the liveness score includes applying, by the computer, the machine-learning architecture to the set of embedding representing the spoofing artifacts to determine the liveness score.

7 . The method of claim 1 , wherein performing the second authentication includes performing, by the computer, the second authentication to restrict the user access to a resource, responsive to the liveness score not satisfying the threshold.

8 . The method of claim 1 , wherein performing the second authentication includes performing, by the computer, the second authentication to permit the user access to a resource, responsive to the liveness score satisfying the threshold.

9 . The method of claim 1 , further comprising generating, by the computer, an indication of a result of the second authentication indicating whether to permit the user access based on the liveness score.

10 . A system for authenticating callers using speech of audio signals in calls, comprising:

a computer comprising one or more processors configured to:

receive, via a first communication channel for a resource, a user authentication request for accessing the resource including authentication credentials for a first authentication;

obtain, via a second communication channel, responsive to successfully authenticating a user according to the first authentication using the authentication credentials, an audio speech signal of the user for a second authentication;

extract a plurality of features from the audio speech signal, including one or more audio background features;

generate a plurality of initial liveness scores using the plurality of features by applying a plurality of machine-learning models of a machine-learning architecture to the plurality of features, wherein each machine-learning model of the plurality of machine-learning models is trained to generate a corresponding initial liveness score of the plurality of initial liveness scores using a corresponding subset of features of the plurality of features extracted from the audio speech signal, the plurality of initial liveness scores including a first initial liveness score indicating an audio background change generated using at least one subset of features having the one or more audio background features from the plurality of features;

generate based upon the plurality of initial liveness scores generated using the plurality of features extracted from the audio speech signal via the second communication channel, a liveness score indicating a likelihood that the user is a human speaker; and

execute the second authentication to determine whether to permit the user access to the resource of the first communication channel based upon comparing the liveness score against a threshold.

11 . The system of claim 10 , wherein, when obtaining the audio speech signal, the computer is further configured to generate a prompt for a user interface, instructing the user to provide the audio speech signal, responsive the computer authenticating the user based on the first authentication of the user using the authentication credentials.

12 . The system of claim 10 , wherein, when generating the liveness score the computer is further configured to apply the machine-learning architecture to the plurality of features to generate the liveness score; and

wherein the machine-learning architecture is trained on a plurality of examples, each example includes a second plurality of features and a label indicating one of human speech or machine-generated speech.

13 . The system of claim 10 , wherein, when generating the plurality of initial liveness scores, the computer is further configured to include at least one of: a second score identifying passive liveness of the speech signal of the speaker, or a third score identifying repetition of speech within the speech signal.

14 . The system of claim 10 wherein, when extracting the plurality of features, the computer is further configured to apply the machine-learning architecture to generate the plurality of features including a set of embeddings representing spoofing artifacts in the audio speech signal.

15 . The system of claim 14 , wherein, when generating the liveness score, the computer is further configured to apply the machine-learning architecture to the set of embedding representing the spoofing artifacts to determine the liveness score.

16 . The system of claim 10 , wherein, when performing the second authentication, the computer is further configured to perform the second authentication to restrict the user access to a resource, responsive to the liveness score not satisfying the threshold.

17 . The system of claim 10 , wherein, when performing the second authentication, the computer is further configured to perform the second authentication to permit the user access to a resource, responsive to the liveness score satisfying the threshold.

18 . The system of claim 10 , wherein the computer is further configured to generate an indication of a result of the second authentication indicating whether to permit the user access based on the liveness score.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: ALTAF, UMAIR; PERI, SAI PRADEEP; PHATELA, LAKSHAY; GUPTA, PAYAS; SUN, YITAO; AFANASEVA, SVETLANA; PATIL, KAILASH; KHOURY, ELIE; MAGNETTA, BRADLEY; BALASUBRAMANIYAN, VIJAY; CHEN, TIANXIANG
To: PINDROP SECURITY, INC.
Reel/Frame 067618/0883 →
Continuity (3)
Provisional Application 63503903 · May 23, 2023
Provisional Application 63497587 · Apr 21, 2023
Related Publication 20240355334A1 · Oct 24, 2024
References Cited (141)
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9318114B2 · Zeljkovic et al. · 2016 [cited by applicant]
US 10565498B1 · Zhiyanov · 2020 [cited by examiner]
US 10607148B1 · Niewczas · 2020 [cited by applicant]
US 10657971B1 · Newstadt et al. · 2020 [cited by applicant]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 10853579B2 · Laxman · 2020 [cited by examiner]
US 10878174B1 · Vontobel · 2020 [cited by examiner]
US 11106690B1 · Dhillon · 2021 [cited by examiner]
US 11151385B2 · Iyer et al. · 2021 [cited by applicant]
US 11417339B1 · Wang et al. · 2022 [cited by applicant]
US 11481434B1 · Venti · 2022 [cited by examiner]
US 11532301B1 · Hajebi · 2022 [cited by examiner]
US 11862177B2 · Chen et al. · 2024 [cited by applicant]
US 11943387B1 · Wolinsky et al. · 2024 [cited by applicant]
US 11972760B1 · Chavez et al. · 2024 [cited by applicant]
US 12026241B2 · Lesso · 2024 [cited by applicant]
US 12062105B2 · Liberman et al. · 2024 [cited by applicant]
US 12206820B2 · Dropuljic et al. · 2025 [cited by applicant]
US 20030115191A1 · Copperman · 2003 [cited by examiner]
US 20050240779A1 · Aull et al. · 2005 [cited by applicant]
US 20080263661A1 · Bouzida · 2008 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by applicant]
US 20090319274A1 · Gross · 2009 [cited by examiner]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20120215776A1 · Guha · 2012 [cited by examiner]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20140270114A1 · Kolbegger et al. · 2014 [cited by applicant]
US 20150278341A1 · Shen · 2015 [cited by examiner]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150347902A1 · Butler et al. · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios et al. · 2016 [cited by applicant]
US 20160323398A1 · Guo · 2016 [cited by examiner]
US 20160328547A1 · Gross · 2016 [cited by applicant]
US 20180077101A1 · Desouza Sana · 2018 [cited by examiner]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180261227A1 · Blouet · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190075123A1 · Smith et al. · 2019 [cited by applicant]
US 20190087870A1 · Gardyne et al. · 2019 [cited by applicant]
US 20190114496A1 · Lesso · 2019 [cited by examiner]
US 20190114497A1 · Lesso · 2019 [cited by examiner]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190206423A1 · Winton · 2019 [cited by examiner]
US 20190311722A1 · Caldwell · 2019 [cited by examiner]
US 20190378520A1 · Chiu · 2019 [cited by applicant]
US 20200044852A1 · Streit · 2020 [cited by examiner]
US 20200118544A1 · Lee · 2020 [cited by examiner]
US 20200145816A1 · Morin et al. · 2020 [cited by applicant]
US 20200153535A1 · Jayaweera Kankanamge et al. · 2020 [cited by applicant]
US 20200184979A1 · Keret et al. · 2020 [cited by applicant]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta et al. · 2020 [cited by applicant]
US 20200314246A1 · Hart · 2020 [cited by applicant]
US 20200344251A1 · Jeyakumar et al. · 2020 [cited by applicant]
US 20200366671A1 · Larson et al. · 2020 [cited by applicant]
US 20210110813A1 · Khoury et al. · 2021 [cited by applicant]
US 20210125619A1 · Lopez Espejo et al. · 2021 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by examiner]
US 20210168238A1 · Adolphe et al. · 2021 [cited by applicant]
US 20210182584A1 · Ionita · 2021 [cited by applicant]
US 20210193174A1 · Enzinger et al. · 2021 [cited by applicant]
US 20210233541A1 · Chen et al. · 2021 [cited by applicant]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20210312906A1 · Kuo et al. · 2021 [cited by applicant]
US 20210326421A1 · Khoury et al. · 2021 [cited by applicant]
US 20210335354A1 · Park · 2021 [cited by applicant]
US 20210406568A1 · Liberman et al. · 2021 [cited by applicant]
US 20220006899A1 · Phatak et al. · 2022 [cited by applicant]
US 20220027648A1 · Trundle et al. · 2022 [cited by applicant]
US 20220059121A1 · Rao et al. · 2022 [cited by applicant]
US 20220060578A1 · Kee et al. · 2022 [cited by applicant]
US 20220076077A1 · Reddy et al. · 2022 [cited by applicant]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220124195A1 · Lu et al. · 2022 [cited by applicant]
US 20220156376A1 · Dos Santos Silva et al. · 2022 [cited by applicant]
US 20220215845A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269761A1 · Steelberg et al. · 2022 [cited by applicant]
US 20220269922A1 · Mathews · 2022 [cited by applicant]
US 20220328050A1 · Hennig · 2022 [cited by examiner]
US 20220366901A1 · Rathaur · 2022 [cited by examiner]
US 20220366916A1 · dos Santos · 2022 [cited by examiner]
US 20230008613A1 · Do et al. · 2023 [cited by applicant]
US 20230032728A1 · Hao · 2023 [cited by applicant]
US 20230082094A1 · Keret et al. · 2023 [cited by applicant]
US 20230107624A1 · Keith · 2023 [cited by applicant]
US 20230134791A1 · Londeree · 2023 [cited by examiner]
US 20230153815A1 · Blouet et al. · 2023 [cited by applicant]
US 20230161853A1 · Nair et al. · 2023 [cited by applicant]
US 20230169370A1 · Elmasry et al. · 2023 [cited by applicant]
US 20230179823A1 · Marten et al. · 2023 [cited by applicant]
US 20230196396A1 · Doumar · 2023 [cited by examiner]
US 20230260521A1 · Slocum et al. · 2023 [cited by applicant]
US 20230262160A1 · Trivedi et al. · 2023 [cited by applicant]
US 20230273978A1 · Suhanic et al. · 2023 [cited by applicant]
US 20230274758A1 · Markhasin et al. · 2023 [cited by applicant]
US 20230308465A1 · Alroobaea et al. · 2023 [cited by applicant]
US 20230386506A1 · Shor et al. · 2023 [cited by applicant]
US 20240040035A1 · Dropuljic et al. · 2024 [cited by applicant]
US 20240040066A1 · Liu et al. · 2024 [cited by applicant]
US 20240073219A1 · Maizels et al. · 2024 [cited by applicant]
US 20240127630A1 · Michaeli et al. · 2024 [cited by applicant]
US 20240127825A1 · Carroll et al. · 2024 [cited by applicant]
US 20240127826A1 · Wolfston et al. · 2024 [cited by applicant]
US 20240203408A1 · Li et al. · 2024 [cited by applicant]
US 20240249712A1 · Yao et al. · 2024 [cited by applicant]
US 20240296698A1 · Matias et al. · 2024 [cited by applicant]
US 20240346850A1 · Kolla et al. · 2024 [cited by applicant]
US 20240363100A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363119A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363123A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363124A1 · Khoury et al. · 2024 [cited by applicant]
US 20240363125A1 · Khoury et al. · 2024 [cited by applicant]
US 20240371510A1 · Aman · 2024 [cited by applicant]
US 20240378274A1 · Tanski · 2024 [cited by examiner]
US 20240406227A1 · Pandit et al. · 2024 [cited by applicant]
US 20240420725A1 · Caven et al. · 2024 [cited by applicant]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by applicant]
US 20250078841A1 · Wu et al. · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar · 2025 [cited by examiner]
US 20250273232A1 · Seo et al. · 2025 [cited by applicant]
EP 3156978A1 · 2017 [cited by applicant]
FR 3105479A1 · 2021 [cited by examiner]
GB 2612032A · 2023 [cited by examiner]
WO WO2008047339A2 · 2008 [cited by examiner]
WO WO2021154600A1 · 2021 [cited by applicant]
WO WO2022082036A1 · 2022 [cited by applicant]
Biometric Authentication System for Industrial Applications using Speaker Recognition (Year: 2020). [cited by examiner]
An Overview of Statistical Pattern Recognition Techniques for Speaker Verification (Year: 2011). [cited by examiner]
Using Deep Learning Neural Networks to Recognize and Authenticate the Identity of the Speaker (Year: 2024). [cited by examiner]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
Mittal et al. “AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US 24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
Shiota et al., “Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification” published in Interspeech 2015 16th Annual Conference of the International Speech Communic… [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]
Final Office Action for U.S. Appl. No. 18/388,428, dated Dec. 18, 2025. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/388,447, dated Dec. 8, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/388,428, dated Apr. 10, 2026. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/646,431, dated Jul. 15, 2026. [cited by applicant]