IP Library Granted Patent US 12,676,144
Granted Patent B2
US 12,676,144 · App. 18/439,049 · Granted Jul 7, 2026

Deepfake detection

Inventors: Umair Altaf (Atlanta, GA); Sai Pradeep Peri (Atlanta, GA); Lakshay Phatela (Atlanta, GA); Payas Gupta (Atlanta, GA); Yitao Sun (Atlanta, GA); Svetlana Afanaseva (Atlanta, GA); Kailash Patil (Atlanta, GA); Elie Khoury (Atlanta, GA); Bradley Magnetta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Tianxiang Chen (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L15/08G06N20/00G10L15/02G10L15/16G10L15/26G10L17/06G10L17/18G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,676,144
App. No.
18/439,049
Filed
Feb 12, 2024
Granted
Jul 7, 2026
Kind
B2
Art Unit
2659
USPC
704/232
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.

Claims (49)

1 . A computer-implemented method of extracting fakeprints to evaluate risk of callers, comprising:

obtaining, by a computer, from a calling device, a raw audio signal including a speech signal for a speaker;

identifying, by the computer, using the raw audio signal, a plurality of datasets in accordance with a corresponding plurality of types of embeddings, the plurality of types of embeddings comprising a textual type of embedding and an audio type of embedding;

extracting, by the computer, executing a fakeprint extractor of a machine-learning architecture, a plurality of fakeprints for the speaker corresponding to the plurality of types of embeddings using the plurality of datasets, the plurality of fakeprints including a first fakeprint corresponding to the textual type of embedding and a second fakeprint corresponding to the audio type of embedding;

determining, by the computer executing a fakeprint evaluator of the machine- learning architecture, a risk score indicating a likelihood that the speaker is fake based on the plurality of fakeprints; and

generating, by the computer, a classification of the speaker as one of human or fake based on the risk score for the speech signal.

2 . The method of claim 1 , further comprising:

identifying, by the computer, a second plurality of fakeprints for the plurality of types of embeddings from a second raw audio signal associated with at least one of the calling device or the speaker;

comparing, by the computer, the first plurality of fakeprints with the second plurality of fakeprints to generate a similarity metric; and

wherein determining the risk score indicating the likelihood that the speaker is fake further comprises determining the risk score based on the similarity metric.

3 . The method of claim 2 , further comprising selecting, by the computer, a plurality of second raw audio signals associated with at least one of the calling device or the speaker received prior to the raw audio signal;

wherein identifying the second plurality of fakeprints further comprises identifying the second plurality of fakeprints for each of the plurality of second raw audio signals.

4 . The method of claim 1 , further comprising generating, by the computer, textual content from the speech signal for the speaker from the raw audio signal; and

wherein extracting the plurality of fakeprints further comprises extracting, using the textual content, at least one fakeprint including the first fakeprint of the plurality of fakeprints corresponding to the textual type of embedding of the plurality of types of embeddings.

5 . The method of claim 1 , further comprising detecting, by the computer, a plurality of temporal segments within the speech signal for the speaker from the raw audio signal, each of the plurality of temporal segments corresponding to a dialogue between the caller and an agent; and

wherein extracting the plurality of fakeprints further comprises extracting, using at least one of the plurality of temporal segments, at least one fakeprint of the plurality of fakeprints corresponding to a temporal type of embedding of the plurality of types of embeddings.

6 . The method of claim 1 , further comprising identifying, by the computer, metadata associated with the raw audio signal, and

extracting the plurality of fakeprints further comprises extracting, using the metadata associated with the raw audio signal, at least one fakeprint of the plurality of fakeprints corresponding to a metadata type of embedding of the plurality of types of embeddings.

7 . The method of claim 1 , further comprising storing, by the computer, an association between the plurality of fakeprints and at least one of the calling device or the speaker, the plurality of fakeprints to be compared against a second plurality of fakeprints extracted from a second raw audio signal from the calling device including a second speech signal of the speaker.

8 . The method of claim 1 , wherein determining the risk score further comprises determining a plurality of risk scores using the plurality of fakeprints, each of the plurality of risk scores corresponding to a respective type of embedding of the plurality of types of embeddings.

9 . The method of claim 1 , wherein obtaining the raw audio signal includes identifying, by the computer, from the raw audio signal, the speech signal corresponding to passive speech of the speaker.

10 . The method of claim 1 , further comprising providing, by the computer, via an interface, an indication of the classification of the speaker as one of human or fake.

11 . A system for extracting fakeprints to evaluate risk of callers, comprising:

a computer comprising one or more processors configured to:

obtain, from a calling device, a raw audio signal including a speech signal for a speaker;

identify, using the raw audio signal, a plurality of datasets in accordance with a corresponding plurality of types of embeddings, the plurality of types of embeddings comprising a textual type of embedding and an audio type of embedding;

extract, executing a fakeprint extractor of a machine-learning architecture, a plurality of fakeprints for the speaker corresponding to the plurality of types of embeddings from the raw audio signal, the plurality of fakeprints including a first fakeprint corresponding to the textual type of embedding and a second fakeprint corresponding to the audio type of embedding;

determine, executing a fakeprint evaluator of the machine-learning architecture, a risk score indicating a likelihood that the speaker is fake based on the plurality of fakeprints; and

generate a classification of the speaker as one of human or fake based on the risk score for the speech signal.

12 . The system of claim 11 , wherein the computer is further configured to:

identify a second plurality of fakeprints for the plurality of types of embeddings from a second raw audio signal associated with at least one of the calling device or the speaker;

compare the first plurality of fakeprints with the second plurality of fakeprints to generate a similarity metric; and

determine the risk score indicating the likelihood that the speaker is fake based on the similarity metric.

13 . The system of claim 12 , wherein the computer is further configured to:

select a plurality of second raw audio signals associated with at least one of the calling device or the speaker received prior to the raw audio signal;

identify the second plurality of fakeprints for each of the plurality of second raw audio signals.

14 . The system of claim 11 , wherein the computer is further configured to:

generate textual content from the speech signal for the speaker from the raw audio signal; and

extract, using the textual content, at least one fakeprint including the first fakeprint of the plurality of fakeprints corresponding to the textual type of embedding of the plurality of types of embeddings.

15 . The system of claim 11 , wherein the computer is further configured to:

detect a plurality of temporal segments within the speech signal for the speaker from the raw audio signal, each of the plurality of temporal segments corresponding to a dialogue between the caller and an agent; and

extract, using at least one of the plurality of temporal segments, at least one fakeprint of the plurality of fakeprints corresponding to a temporal type of embedding of the plurality of types of embeddings.

16 . The system of claim 11 , wherein the computer is further configured to:

identify metadata associated with the raw audio signal, and

extract, using the metadata associated with the raw audio signal, at least one fakeprint of the plurality of fakeprints corresponding to a metadata type of embedding of the plurality of types of embeddings.

17 . The system of claim 11 , wherein the computer is further configured to store an association between the plurality of fakeprints and at least one of the calling device or the speaker, the plurality of fakeprints to be compared against a second plurality of fakeprints extracted from a second raw audio signal from the calling device including a second speech signal of the speaker.

18 . The system of claim 11 , wherein the computer is further configured to determine a plurality of risk scores using the plurality of fakeprints, each of the plurality of risk scores corresponding to a respective type of embedding of the plurality of types of embeddings.

19 . The system of claim 11 , wherein the computer is further configured to identify, from the raw audio signal, the speech signal corresponding to passive speech of the speaker.

20 . The system of claim 11 , wherein the computer is further configured to provide, via an interface, an indication of the classification of the speaker as one of human or fake.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2024
From: ALTAF, UMAIR; PERI, SAI PRADEEP; PHATELA, LAKSHAY; GUPTA, PAYAS; SUN, YITAO; AFANASEVA, SVETLANA; PATIL, KAILASH; KHOURY, ELIE; MAGNETTA, BRADLEY; BALASUBRAMANIYAN, VIJAY; CHEN, TIANXIANG
To: PINDROP SECURITY, INC.
Reel/Frame 066444/0238 →
Continuity (3)
Provisional Application 63503903 · May 23, 2023
Provisional Application 63497587 · Apr 21, 2023
Related Publication 20240355336A1 · Oct 24, 2024
References Cited (130)
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9318114B2 · Zeljkovic et al. · 2016 [cited by applicant]
US 10565498B1 · Zhiyanov · 2020 [cited by applicant]
US 10607148B1 · Niewczas · 2020 [cited by examiner]
US 10657971B1 · Newstadt et al. · 2020 [cited by applicant]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 10853579B2 · Laxman et al. · 2020 [cited by applicant]
US 10878174B1 · Vontobel et al. · 2020 [cited by applicant]
US 11106690B1 · Dhillon et al. · 2021 [cited by applicant]
US 11151385B2 · Iyer et al. · 2021 [cited by applicant]
US 11417339B1 · Wang · 2022 [cited by examiner]
US 11481434B1 · Venti et al. · 2022 [cited by applicant]
US 11532301B1 · Hajebi et al. · 2022 [cited by applicant]
US 11943387B1 · Wolinsky et al. · 2024 [cited by applicant]
US 11972760B1 · Chavez · 2024 [cited by examiner]
US 12026241B2 · Lesso · 2024 [cited by applicant]
US 12062105B2 · Liberman · 2024 [cited by examiner]
US 12087319B1 · Looney et al. · 2024 [cited by applicant]
US 12206820B2 · Dropuljic et al. · 2025 [cited by applicant]
US 20030115191A1 · Copperman et al. · 2003 [cited by applicant]
US 20050240779A1 · Aull et al. · 2005 [cited by applicant]
US 20080263661A1 · Bouzida · 2008 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by applicant]
US 20090319274A1 · Gross · 2009 [cited by applicant]
US 20100131273A1 · Aley-Raz et al. · 2010 [cited by applicant]
US 20120215776A1 · Guha et al. · 2012 [cited by applicant]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20140270114A1 · Kolbegger et al. · 2014 [cited by applicant]
US 20150278341A1 · Shen · 2015 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150347902A1 · Butler et al. · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios et al. · 2016 [cited by applicant]
US 20160323398A1 · Guo et al. · 2016 [cited by applicant]
US 20160328547A1 · Gross · 2016 [cited by applicant]
US 20180077101A1 · Desouza Sana et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180261227A1 · Blouet · 2018 [cited by applicant]
US 20190037081A1 · Rao et al. · 2019 [cited by applicant]
US 20190075123A1 · Smith et al. · 2019 [cited by applicant]
US 20190087870A1 · Gardyne et al. · 2019 [cited by applicant]
US 20190114496A1 · Lesso · 2019 [cited by applicant]
US 20190114497A1 · Lesso · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190206423A1 · Winton et al. · 2019 [cited by applicant]
US 20190311722A1 · Caldwell · 2019 [cited by applicant]
US 20190378520A1 · Chiu · 2019 [cited by applicant]
US 20200044852A1 · Streit · 2020 [cited by applicant]
US 20200118544A1 · Lee et al. · 2020 [cited by applicant]
US 20200145816A1 · Morin et al. · 2020 [cited by applicant]
US 20200153535A1 · Jayaweera Kankanamge et al. · 2020 [cited by applicant]
US 20200184979A1 · Keret et al. · 2020 [cited by applicant]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta et al. · 2020 [cited by applicant]
US 20200314246A1 · Hart · 2020 [cited by applicant]
US 20200344251A1 · Jeyakumar et al. · 2020 [cited by applicant]
US 20200366671A1 · Larson · 2020 [cited by examiner]
US 20210125619A1 · Lopez Espejo et al. · 2021 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by applicant]
US 20210168238A1 · Adolphe et al. · 2021 [cited by applicant]
US 20210182584A1 · Ionita · 2021 [cited by applicant]
US 20210193174A1 · Enzinger et al. · 2021 [cited by applicant]
US 20210233541A1 · Chen et al. · 2021 [cited by applicant]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20210312906A1 · Kuo et al. · 2021 [cited by applicant]
US 20210335354A1 · Park · 2021 [cited by applicant]
US 20210406568A1 · Liberman et al. · 2021 [cited by applicant]
US 20220006899A1 · Phatak et al. · 2022 [cited by applicant]
US 20220027648A1 · Trundle et al. · 2022 [cited by applicant]
US 20220059121A1 · Rao et al. · 2022 [cited by applicant]
US 20220060578A1 · Kee et al. · 2022 [cited by applicant]
US 20220076077A1 · Reddy et al. · 2022 [cited by applicant]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220108701A1 · Gupta et al. · 2022 [cited by applicant]
US 20220124195A1 · Lu et al. · 2022 [cited by applicant]
US 20220156376A1 · Dos Santos Silva et al. · 2022 [cited by applicant]
US 20220215845A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269761A1 · Steelberg · 2022 [cited by examiner]
US 20220269922A1 · Mathews · 2022 [cited by examiner]
US 20220328050A1 · Hennig et al. · 2022 [cited by applicant]
US 20220366901A1 · Rathaur et al. · 2022 [cited by applicant]
US 20220366916A1 · Dos Santos et al. · 2022 [cited by applicant]
US 20230008613A1 · Do et al. · 2023 [cited by applicant]
US 20230032728A1 · Hao · 2023 [cited by applicant]
US 20230082094A1 · Keret et al. · 2023 [cited by applicant]
US 20230107624A1 · Keith · 2023 [cited by applicant]
US 20230134791A1 · Londeree · 2023 [cited by applicant]
US 20230153815A1 · Blouet et al. · 2023 [cited by applicant]
US 20230161853A1 · Nair et al. · 2023 [cited by applicant]
US 20230169370A1 · Elmasry et al. · 2023 [cited by applicant]
US 20230179823A1 · Marten et al. · 2023 [cited by applicant]
US 20230196396A1 · Doumar et al. · 2023 [cited by applicant]
US 20230260521A1 · Slocum · 2023 [cited by examiner]
US 20230262160A1 · Trivedi et al. · 2023 [cited by applicant]
US 20230273978A1 · Suhanic · 2023 [cited by examiner]
US 20230274758A1 · Markhasin et al. · 2023 [cited by applicant]
US 20230308465A1 · Alroobaea et al. · 2023 [cited by applicant]
US 20230386506A1 · Shor et al. · 2023 [cited by applicant]
US 20240040035A1 · Dropuljic et al. · 2024 [cited by applicant]
US 20240040066A1 · Liu et al. · 2024 [cited by applicant]
US 20240073219A1 · Maizels et al. · 2024 [cited by applicant]
US 20240127630A1 · Michaeli · 2024 [cited by examiner]
US 20240127825A1 · Carroll · 2024 [cited by examiner]
US 20240127826A1 · Wolfston et al. · 2024 [cited by applicant]
US 20240203408A1 · Li et al. · 2024 [cited by applicant]
US 20240249712A1 · Yao et al. · 2024 [cited by applicant]
US 20240296698A1 · Matias et al. · 2024 [cited by applicant]
US 20240346850A1 · Kolla et al. · 2024 [cited by applicant]
US 20240371510A1 · Aman · 2024 [cited by applicant]
US 20240378274A1 · Tanski et al. · 2024 [cited by applicant]
US 20240406227A1 · Pandit et al. · 2024 [cited by applicant]
US 20240420725A1 · Caven et al. · 2024 [cited by applicant]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by examiner]
US 20250078841A1 · Wu et al. · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar et al. · 2025 [cited by applicant]
US 20250273232A1 · Seo et al. · 2025 [cited by applicant]
EP 3156978A1 · 2017 [cited by applicant]
FR 3105479A1 · 2021 [cited by applicant]
GB 2612032A · 2023 [cited by applicant]
WO WO2008047339A2 · 2008 [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
Mittal et al. “AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]
Shiota et al., “Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification” published in Interspeech 2015 16th Annual Conference of the International Speech Communic… [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]
Final Office Action for U.S. Appl. No. 18/388,364, dated Mar. 9, 2026. [cited by applicant]
Al-Rahman et al., “Using Deep Learning Neural Networks to Recognize and Authenticate the Identity of the Speaker”, 2024 21st International Multi-Conference on Systems, Signals & Devices (SSD). (8 pages). [cited by applicant]
Fazel et al., “An Overview of Statistical Pattern Recognition Techniques for Speaker Verification”, IEEE Circuits and Systems Magazine, Second Quarter 2011. (20 pages). [cited by applicant]
Shayamunda et al., “Biometric Authentication System for Industrial Applications using Speaker Recognition”, Department of Electrical, Electronic and Computer Engineering, University of Pretoria Tshwane, South Africa, 20… [cited by applicant]