IP Library Granted Patent US 12,562,150
Granted Patent B2
US 12,562,150 · App. 18/388,466 · Granted Feb 24, 2026

Deepfake detection

Inventors: Umair Altaf (Atlanta, GA); Sai Pradeep Peri (Atlanta, GA); Lakshay Phatela (Atlanta, GA); Payas Gupta (Atlanta, GA); Yitao Sun (Atlanta, GA); Svetlana Afanaseva (Atlanta, GA); Kailash Patil (Atlanta, GA); Elie Khoury (Atlanta, GA); Bradley Magnetta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Tianxiang Chen (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L15/02G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,150
App. No.
18/388,466
Granted
Feb 24, 2026
Kind
B2
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.

Claims (40)

1 . A computer-implemented method for detecting machine-based speech in calls, the method comprising:

in response to detecting a first speech region in a first audio signal received during a call from a caller, applying, by a computer, an event classifier on a first set of acoustic features extracted from the first audio signal to classify a first set of background audio events in the first audio signal;

in response to detecting a second speech region in a second audio signal received during the call, applying, by the computer, the event classifier on a second set of acoustic features extracted from the second audio signal to classify a second set of background audio events in the second audio signal;

generating, by the computer, a machine likelihood score based upon an amount of similarity between the first set of background audio events and the second set of background audio events obtained during the call; and

identifying, by the computer, the caller as a machine caller in response to determining that the machine likelihood score satisfies a similarity threshold.

2 . The method of claim 1 , wherein generating the machine likelihood score further comprises generating the machine likelihood score based upon identifying a first background event in the first audio signal and a second background event in the second audio signal within a time period during the call, whereby the machine caller provides a plurality of sets of background audio events that are distinct throughout the call.

3 . The method of claim 1 , wherein generating the machine likelihood score further comprises applying an attack prediction model to the first set of background events and the second set of background events to generate the machine likelihood score over a first segment corresponding to the first audio signal and a second segment corresponding to the second audio signal.

4 . The method of claim 1 , wherein applying the event classifier on the first set of acoustic features includes applying, by the computer, the event classifier to detect one or more first anomalous events within the first audio signal; and

wherein applying the event classifier on the second set of acoustic features includes applying, by the computer, the event classifier to detect one or more second anomalous events within the second audio signal.

5 . The method of claim 1 , further comprising:

applying, by the computer, an estimation model to the first set of acoustic features to determine a first quality metric indicating a first degree of change in quality through the first audio signal; and

applying, by the computer, the estimation model to the second set of acoustic features to determine a second quality metric indicating a second degree of change in quality through the second audio signal.

6 . The method of claim 5 , wherein the first quality metric and the second quality metric each include at least one of: (i) a respective noise type, (ii) a respective reverberation ratio, or (iii) a respective signal to noise (SNR) ratio.

7 . The method of claim 5 , wherein generating the machine likelihood score includes generating, by the computer, the machine likelihood score based upon a comparison between the first quality metric and the second quality metric.

8 . The method of claim 1 , further comprising:

extracting, by the computer, the first set of acoustic features from a first background in the first audio signal; and

extracting, by the computer, the second set of acoustic features from a second background in the second audio signal.

9 . The method of claim 1 , further comprising identifying, by the computer, the caller as a human caller in response to determining that the machine likelihood score does not satisfy the similarity threshold.

10 . The method of claim 1 , further comprising generating, by the computer, an indication, for a user interface, of the caller as a machine caller in response to determining that the machine likelihood score satisfies the threshold.

11 . A system for detecting machine-based speech in calls, comprising:

a computer comprising one or more processors and configured to:

detect a first speech region in a first audio signal received during a call from a caller;

in response to the computer detecting the first speech region in the first audio signal received during the call from the caller, apply an event classifier on a first set of acoustic features extracted from the first audio signal to classify a first set of background audio events in the first audio signal;

in response to detecting a second speech region in a second audio signal received during the call, apply the event classifier on a second set of acoustic features extracted from the second audio signal to classify a second set of background audio events in the second audio signal;

generate a machine likelihood score based upon an amount of similarity between the first set of background audio events and the second set of background audio events obtained during the call; and

identify the caller as a machine caller in response to determining that the machine likelihood score satisfies a similarity threshold.

12 . The system of claim 11 , wherein, when generating the machine likelihood score, the computer is further configured to generate the machine likelihood score based upon identifying a first background event in the first audio signal and a second background event in the second audio signal within a time period during the call, whereby the machine caller provides a plurality of sets of background audio events that are distinct throughout the call.

13 . The system of claim 11 , wherein, when generating the machine likelihood score, the computer is further configured to apply an attack prediction model to the first set of background events and the second set of background events to generate the machine likelihood score over a first segment corresponding to the first audio signal and a second segment corresponding to the second audio signal.

14 . The system of claim 11 , wherein, when applying the event classifier on the first set of acoustic features, the computer is further configured to apply the event classifier to detect one or more first anomalous events within the first audio signal; and

wherein, when applying the event classifier on the second set of acoustic features, the computer is further configured to apply the event classifier to detect one or more second anomalous events within the second audio signal.

15 . The system of claim 11 , wherein the computer is further configured to:

apply an estimation model to the first set of acoustic features to determine a first quality metric indicating a first degree of change in quality through the first audio signal; and

apply the estimation model to the second set of acoustic features to determine a second quality metric indicating a second degree of change in quality through the second audio signal.

16 . The system of claim 15 , wherein the first quality metric and the second quality metric each include at least one of: (i) a respective noise type, (ii) a respective reverberation ratio, or (iii) a respective signal to noise (SNR) ratio.

17 . The system of claim 15 , wherein, when generating the machine likelihood score, the computer is further configured to generate the machine likelihood score based upon a comparison between the first quality metric and the second quality metric.

18 . The system of claim 11 , wherein the computer is further configured to:

extract the first set of acoustic features from a first background in the first audio signal; and

extract the second set of acoustic features from a second background in the second audio signal.

19 . The system of claim 11 , wherein the computer is further configured to identify the caller as a human caller in response to the computer determining that the machine likelihood score does not satisfy the similarity threshold.

20 . The system of claim 11 , wherein the computer is further configured to generate an indication, for a user interface, of the caller as a machine caller in response to determining that the machine likelihood score satisfies the threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: ALTAF, UMAIR; PERI, SAI PRADEEP; PHATELA, LAKSHAY; GUPTA, PAYAS; SUN, YITAO; AFANASEVA, SVETLANA; PATIL, KAILASH; KHOURY, ELIE; MAGNETTA, BRADLEY; BALASUBRAMANIYAN, VIJAY; CHEN, TIANXIANG
To: PINDROP SECURITY, INC.
Reel/Frame 067618/0883 →
Continuity (3)
Provisional Application 63503903 · May 23, 2023
Provisional Application 63497587 · Apr 21, 2023
Related Publication 20240363099A1 · Oct 31, 2024
References Cited (97)
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9318114B2 · Zeljkovic et al. · 2016 [cited by applicant]
US 10565498B1 · Zhiyanov · 2020 [cited by applicant]
US 10607148B1 · Niewczas · 2020 [cited by applicant]
US 10657971B1 · Newstadt et al. · 2020 [cited by applicant]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 10853579B2 · Laxman et al. · 2020 [cited by applicant]
US 10878174B1 · Vontobel et al. · 2020 [cited by applicant]
US 11106690B1 · Dhillon et al. · 2021 [cited by applicant]
US 11151385B2 · Iyer et al. · 2021 [cited by applicant]
US 11417339B1 · Wang et al. · 2022 [cited by applicant]
US 11481434B1 · Venti et al. · 2022 [cited by applicant]
US 11532301B1 · Hajebi et al. · 2022 [cited by applicant]
US 11972760B1 · Chavez et al. · 2024 [cited by applicant]
US 12062105B2 · Liberman et al. · 2024 [cited by applicant]
US 20030115191A1 · Copperman et al. · 2003 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by applicant]
US 20090319274A1 · Gross · 2009 [cited by applicant]
US 20120215776A1 · Guha et al. · 2012 [cited by applicant]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20140270114A1 · Kolbegger et al. · 2014 [cited by applicant]
US 20150278341A1 · Shen · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios et al. · 2016 [cited by applicant]
US 20160323398A1 · Guo et al. · 2016 [cited by applicant]
US 20160328547A1 · Gross · 2016 [cited by applicant]
US 20180077101A1 · Desouza Sana et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180261227A1 · Blouet · 2018 [cited by applicant]
US 20190087870A1 · Gardyne et al. · 2019 [cited by applicant]
US 20190114496A1 · Lesso · 2019 [cited by applicant]
US 20190114497A1 · Lesso · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190206423A1 · Winton et al. · 2019 [cited by applicant]
US 20190311722A1 · Caldwell · 2019 [cited by applicant]
US 20190378520A1 · Chiu · 2019 [cited by applicant]
US 20200044852A1 · Streit · 2020 [cited by applicant]
US 20200118544A1 · Lee et al. · 2020 [cited by applicant]
US 20200184979A1 · Keret et al. · 2020 [cited by applicant]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta et al. · 2020 [cited by applicant]
US 20200314246A1 · Hart · 2020 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by applicant]
US 20210168238A1 · Adolphe et al. · 2021 [cited by applicant]
US 20210182584A1 · Ionita · 2021 [cited by applicant]
US 20210193174A1 · Enzinger et al. · 2021 [cited by applicant]
US 20210233541A1 · Chen · 2021 [cited by examiner]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20210312906A1 · Kuo et al. · 2021 [cited by applicant]
US 20210335354A1 · Park · 2021 [cited by applicant]
US 20220027648A1 · Trundle et al. · 2022 [cited by applicant]
US 20220060578A1 · Kee et al. · 2022 [cited by applicant]
US 20220076077A1 · Reddy · 2022 [cited by examiner]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220124195A1 · Lu et al. · 2022 [cited by applicant]
US 20220215845A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269922A1 · Mathews · 2022 [cited by applicant]
US 20220328050A1 · Hennig et al. · 2022 [cited by applicant]
US 20220366901A1 · Rathaur et al. · 2022 [cited by applicant]
US 20220366916A1 · Dos Santos et al. · 2022 [cited by applicant]
US 20230008613A1 · Do et al. · 2023 [cited by applicant]
US 20230032728A1 · Hao · 2023 [cited by applicant]
US 20230082094A1 · Keret et al. · 2023 [cited by applicant]
US 20230107624A1 · Keith · 2023 [cited by applicant]
US 20230134791A1 · Londeree · 2023 [cited by applicant]
US 20230153815A1 · Blouet et al. · 2023 [cited by applicant]
US 20230161853A1 · Nair et al. · 2023 [cited by applicant]
US 20230179823A1 · Marten et al. · 2023 [cited by applicant]
US 20230196396A1 · Doumar et al. · 2023 [cited by applicant]
US 20230260521A1 · Slocum et al. · 2023 [cited by applicant]
US 20230262160A1 · Trivedi et al. · 2023 [cited by applicant]
US 20230274758A1 · Markhasin et al. · 2023 [cited by applicant]
US 20230386506A1 · Shor et al. · 2023 [cited by applicant]
US 20240040066A1 · Liu et al. · 2024 [cited by applicant]
US 20240073219A1 · Maizels et al. · 2024 [cited by applicant]
US 20240127630A1 · Michaeli et al. · 2024 [cited by applicant]
US 20240203408A1 · Li et al. · 2024 [cited by applicant]
US 20240249712A1 · Yao et al. · 2024 [cited by applicant]
US 20240296698A1 · Matias et al. · 2024 [cited by applicant]
US 20240346850A1 · Kolla et al. · 2024 [cited by applicant]
US 20240371510A1 · Aman · 2024 [cited by applicant]
US 20240378274A1 · Tanski et al. · 2024 [cited by applicant]
US 20240406227A1 · Pandit et al. · 2024 [cited by applicant]
US 20240420725A1 · Caven et al. · 2024 [cited by applicant]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar et al. · 2025 [cited by applicant]
US 20250273232A1 · Seo et al. · 2025 [cited by applicant]
EP 3156978A1 · 2017 [cited by applicant]
FR 3105479A1 · 2021 [cited by applicant]
GB 2612032A · 2023 [cited by applicant]
WO WO2008047339A2 · 2008 [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US 24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
Mittal et al. “AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]
Shiota et al., “Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification” published in INTERSPEECH 2015 16th Annual Conference of the International Speech Communic… [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]