IP Library Granted Patent US 12,592,220
Granted Patent B2
US 12,592,220 · App. 18/388,447 · Granted Mar 31, 2026

Deepfake detection

Inventors: Umair Altaf (Atlanta, GA); Sai Pradeep Peri (Atlanta, GA); Lakshay Phatela (Atlanta, GA); Payas Gupta (Atlanta, GA); Yitao Sun (Atlanta, GA); Svetlana Afanaseva (Atlanta, GA); Kailash Patil (Atlanta, GA); Elie Khoury (Atlanta, GA); Bradley Magnetta (Atlanta, GA); Vijay Balasubramaniyan (Atlanta, GA); Tianxiang Chen (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L15/08G06N20/00G10L15/02G10L15/16G10L15/26G10L17/06G10L17/18G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,220
App. No.
18/388,447
Granted
Mar 31, 2026
Kind
B2
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.

Claims (40)

1 . A computer-implemented method for generating liveness scores for audio in calls, comprising:

obtaining, by a computer, a raw audio signal from a calling device including a speech signal for a speaker;

identifying, by the computer, a plurality of non-speech segments and a plurality of speech segments in the raw audio signal by executing a voice activity detection engine for detecting instances speech or non-speech in the raw audio signal;

extracting, by the computer, a plurality of background acoustic features for the plurality of non-speech segments and a plurality of acoustic features for the plurality of speech segments by executing one or more neural network architecture on the plurality of non-speech segments and the plurality of speech segments identified in the raw audio signal;

determining, by the computer, a plurality of scores based on the raw audio signal, the plurality of scores comprising (i) a first score identifying background change generated based upon comparing the plurality of background acoustic features extracted for the plurality of non-speech segments in the raw audio signal, and (ii) a second score identifying passive liveness generated based upon the plurality of acoustic features extracted for the plurality of speech segments of the speech signal of the speaker;

applying, by the computer, a machine-learning architecture to the plurality of scores to generate a liveness score indicating a likelihood that the speaker is a human; and

generating, by the computer, a classification of the speaker as one of human or machine based on the liveness score for the speech signal.

2 . The method of claim 1 , further comprising:

receiving, by the computer, a feedback dataset identifying a second classification indicating the speaker as one of human or machine; and

re-training, by the computer, the machine-learning architecture using a comparison of the classification and the second classification of the feedback dataset.

3 . The method of claim 1 , wherein obtaining the raw audio signal includes:

obtaining, by the computer, a training dataset including (i) the raw audio signal and (ii) a label indicating one of machine or human for the raw audio signal; and

updating, by the computer, the machine-learning architecture based on a comparison between the classification from the machine-learning architecture and the label of the training dataset.

4 . The method of claim 1 , further comprising determining, by the computer, based on at least one of feedback or training data, a threshold to compare against the liveness score to generate the classification.

5 . The method of claim 1 , wherein determining the plurality of scores further comprises applying, by the computer, the machine-learning architecture to extract the plurality of features from the raw audio signal and to generate the plurality of scores based on the plurality of features.

6 . The method of claim 1 , wherein determining the plurality of scores includes determining, by the computer, the plurality of scores including a third score identifying repetition of speech within the speech signal.

7 . The method of claim 1 , wherein applying the machine-learning architecture includes applying, by the computer, a plurality of weights defined by the machine-learning architecture to the plurality of scores to generate the liveness score.

8 . The method of claim 1 , wherein applying the machine-learning architecture includes applying, by the computer, a scoring neural network of the machine-learning architecture to generate the liveness score.

9 . The method of claim 1 , wherein obtaining the raw audio signal includes identifying, by the computer, from the raw audio signal, the speech signal corresponding to passive speech of the speaker.

10 . The method of claim 1 , further comprising providing, by the computer, via an interface, an indication of the classification of the speaker as one of human or machine.

11 . A system for generating liveness scores for audio in calls, comprising:

a computer comprising one or more processors and configured to:

obtain, a raw audio signal from a calling device including a speech signal for a speaker;

identify a plurality of non-speech segments and a plurality of speech segments in the raw audio signal by executing a voice activity detection engine for detecting instances speech or non-speech in the raw audio signal;

extract a plurality of background acoustic features for the plurality of non-speech segments and a plurality of acoustic features for the plurality of speech segments by executing one or more neural network architecture on the plurality of non-speech segments and the plurality of speech segments identified in the raw audio signal;

determine a plurality of scores based on the raw audio signal, the plurality of scores comprising (i) a first score identifying background change generated based upon comparing the plurality of background acoustic features extracted for the plurality of non-speech segments in the raw audio signal according to a voice activity detection engine, and (ii) a second score identifying passive liveness generated based upon the plurality of acoustic features extracted for the plurality of speech segments of the speech signal of the speaker;

apply a machine-learning architecture to the plurality of scores to generate a liveness score indicating a likelihood that the speaker is a human; and

generate a classification of the speaker as one of human or machine based on the liveness score for the speech signal.

12 . The system of claim 11 , wherein the computer is further configured to:

receive a feedback dataset identifying a second classification indicating the speaker as one of human or machine; and

re-train the machine-learning architecture using a comparison of the classification and the second classification of the feedback dataset.

13 . The system of claim 11 , wherein, when obtaining the raw audio signal, the computer is further configured to obtain a training dataset including (i) the raw audio signal and (ii) a label indicating one of machine or human for the raw audio signal; and

wherein the computer is further configured to update the machine-learning architecture based on a comparison between the classification from the machine-learning architecture and the label of the training dataset.

14 . The system of claim 11 , wherein the computer is further configured to determine based on at least one of feedback or training data, a threshold to compare against the liveness score to generate the classification.

15 . The system of claim 11 , wherein, when determining the plurality of scores, the computer is further configured to apply the machine-learning architecture to extract the plurality of features from the raw audio signal and to generate the plurality of scores based on the plurality of features.

16 . The system of claim 11 , wherein, where determining the plurality of scores, the computer is further configured to determine the plurality of scores including a third score identifying repetition of speech within the speech signal.

17 . The system of claim 11 , wherein, when applying the machine-learning architecture, the computer is configured to apply a plurality of weights defined by the machine-learning architecture to the plurality of scores to generate the liveness score.

18 . The system of claim 11 , wherein, when applying the machine-learning architecture, the computer is further configured to apply a scoring neural network of the machine-learning architecture to generate the liveness score.

19 . The system of claim 11 , wherein, when obtaining the raw audio signal, the computer is further configured to identify, from the raw audio signal, the speech signal corresponding to passive speech of the speaker.

20 . The system of claim 11 , wherein the computer is further configured to provide via an interface, an indication of the classification of the speaker as one of human or machine.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: ALTAF, UMAIR; PERI, SAI PRADEEP; PHATELA, LAKSHAY; GUPTA, PAYAS; SUN, YITAO; AFANASEVA, SVETLANA; PATIL, KAILASH; KHOURY, ELIE; MAGNETTA, BRADLEY; BALASUBRAMANIYAN, VIJAY; CHEN, TIANXIANG
To: PINDROP SECURITY, INC.
Reel/Frame 067618/0883 →
Continuity (3)
Provisional Application 63497587 · Apr 21, 2023
Provisional Application 63503903 · May 23, 2023
Related Publication 20240355323A1 · Oct 24, 2024
References Cited (107)
US 9060057B1 · Danis · 2015 [cited by applicant]
US 9318114B2 · Zeljkovic et al. · 2016 [cited by applicant]
US 10565498B1 · Zhiyanov · 2020 [cited by applicant]
US 10607148B1 · Niewczas · 2020 [cited by applicant]
US 10657971B1 · Newstadt et al. · 2020 [cited by applicant]
US 10720165B2 · Guo et al. · 2020 [cited by applicant]
US 10853579B2 · Laxman et al. · 2020 [cited by applicant]
US 10878174B1 · Vontobel et al. · 2020 [cited by applicant]
US 11106690B1 · Dhillon et al. · 2021 [cited by applicant]
US 11151385B2 · Iyer et al. · 2021 [cited by applicant]
US 11417339B1 · Wang et al. · 2022 [cited by applicant]
US 11481434B1 · Venti et al. · 2022 [cited by applicant]
US 11532301B1 · Hajebi et al. · 2022 [cited by applicant]
US 11943387B1 · Wolinsky et al. · 2024 [cited by applicant]
US 11972760B1 · Chavez et al. · 2024 [cited by applicant]
US 12026241B2 · Lesso · 2024 [cited by applicant]
US 12062105B2 · Liberman et al. · 2024 [cited by applicant]
US 12087319B1 · Looney et al. · 2024 [cited by applicant]
US 20030115191A1 · Copperman et al. · 2003 [cited by applicant]
US 20050240779A1 · Aull et al. · 2005 [cited by applicant]
US 20090319271A1 · Gross · 2009 [cited by applicant]
US 20090319274A1 · Gross · 2009 [cited by examiner]
US 20120215776A1 · Guha et al. · 2012 [cited by applicant]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20140270114A1 · Kolbegger et al. · 2014 [cited by applicant]
US 20150278341A1 · Shen · 2015 [cited by applicant]
US 20160253548A1 · Dos Remedios et al. · 2016 [cited by applicant]
US 20160323398A1 · Guo et al. · 2016 [cited by applicant]
US 20160328547A1 · Gross · 2016 [cited by applicant]
US 20180077101A1 · Desouza Sana et al. · 2018 [cited by applicant]
US 20180254046A1 · Khoury et al. · 2018 [cited by applicant]
US 20180261227A1 · Blouet · 2018 [cited by applicant]
US 20190087870A1 · Gardyne et al. · 2019 [cited by applicant]
US 20190114496A1 · Lesso · 2019 [cited by applicant]
US 20190114497A1 · Lesso · 2019 [cited by applicant]
US 20190174000A1 · Bharrat et al. · 2019 [cited by applicant]
US 20190206423A1 · Winton et al. · 2019 [cited by applicant]
US 20190311722A1 · Caldwell · 2019 [cited by applicant]
US 20190378520A1 · Chiu · 2019 [cited by applicant]
US 20200044852A1 · Streit · 2020 [cited by applicant]
US 20200118544A1 · Lee · 2020 [cited by examiner]
US 20200184979A1 · Keret et al. · 2020 [cited by applicant]
US 20200194032A1 · Basye et al. · 2020 [cited by applicant]
US 20200311738A1 · Gupta et al. · 2020 [cited by applicant]
US 20200314246A1 · Hart · 2020 [cited by applicant]
US 20210125619A1 · Espejo et al. · 2021 [cited by applicant]
US 20210141896A1 · Streit · 2021 [cited by applicant]
US 20210168238A1 · Adolphe et al. · 2021 [cited by applicant]
US 20210182584A1 · Ionita · 2021 [cited by examiner]
US 20210193174A1 · Enzinger et al. · 2021 [cited by applicant]
US 20210233541A1 · Chen et al. · 2021 [cited by applicant]
US 20210280171A1 · Phatak et al. · 2021 [cited by applicant]
US 20210312906A1 · Kuo et al. · 2021 [cited by applicant]
US 20210335354A1 · Park · 2021 [cited by applicant]
US 20210406568A1 · Liberman et al. · 2021 [cited by applicant]
US 20220027648A1 · Trundle et al. · 2022 [cited by applicant]
US 20220060578A1 · Kee et al. · 2022 [cited by applicant]
US 20220076077A1 · Reddy et al. · 2022 [cited by applicant]
US 20220084509A1 · Sivaraman et al. · 2022 [cited by applicant]
US 20220108701A1 · Gupta et al. · 2022 [cited by applicant]
US 20220124195A1 · Lu et al. · 2022 [cited by applicant]
US 20220215845A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269922A1 · Mathews · 2022 [cited by applicant]
US 20220328050A1 · Hennig et al. · 2022 [cited by applicant]
US 20220366901A1 · Rathaur et al. · 2022 [cited by applicant]
US 20220366916A1 · Dos Santos et al. · 2022 [cited by applicant]
US 20230008613A1 · Do et al. · 2023 [cited by applicant]
US 20230032728A1 · Hao · 2023 [cited by applicant]
US 20230082094A1 · Keret et al. · 2023 [cited by applicant]
US 20230107624A1 · Keith, Jr. · 2023 [cited by examiner]
US 20230134791A1 · Londeree · 2023 [cited by applicant]
US 20230153815A1 · Blouet · 2023 [cited by examiner]
US 20230161853A1 · Nair et al. · 2023 [cited by applicant]
US 20230179823A1 · Marten et al. · 2023 [cited by applicant]
US 20230196396A1 · Doumar et al. · 2023 [cited by applicant]
US 20230260521A1 · Slocum et al. · 2023 [cited by applicant]
US 20230262160A1 · Trivedi et al. · 2023 [cited by applicant]
US 20230274758A1 · Markhasin et al. · 2023 [cited by applicant]
US 20230386506A1 · Shor et al. · 2023 [cited by applicant]
US 20240040035A1 · Dropuljic et al. · 2024 [cited by applicant]
US 20240040066A1 · Liu et al. · 2024 [cited by applicant]
US 20240073219A1 · Maizels et al. · 2024 [cited by applicant]
US 20240127630A1 · Michaeli et al. · 2024 [cited by applicant]
US 20240127826A1 · Wolfston et al. · 2024 [cited by applicant]
US 20240203408A1 · Li et al. · 2024 [cited by applicant]
US 20240249712A1 · Yao et al. · 2024 [cited by applicant]
US 20240296698A1 · Matias et al. · 2024 [cited by applicant]
US 20240346850A1 · Kolla et al. · 2024 [cited by applicant]
US 20240371510A1 · Aman · 2024 [cited by applicant]
US 20240378274A1 · Tanski et al. · 2024 [cited by applicant]
US 20240406227A1 · Pandit et al. · 2024 [cited by applicant]
US 20240420725A1 · Caven et al. · 2024 [cited by applicant]
US 20250022473A1 · Eilam Tzoreff · 2025 [cited by applicant]
US 20250078841A1 · Wu et al. · 2025 [cited by applicant]
US 20250086213A1 · Dilipkumar et al. · 2025 [cited by applicant]
US 20250273232A1 · Seo et al. · 2025 [cited by applicant]
EP 3156978A1 · 2017 [cited by examiner]
FR 3105479A1 · 2021 [cited by applicant]
GB 2612032A · 2023 [cited by applicant]
WO WO2008047339A2 · 2008 [cited by applicant]
Shiota et al. 2015, Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification. in INTERSPEECH 2015 16th Annual Conference of the International Speech Communication … [cited by examiner]
International Search Report and Written Opinion issued in International Application No. PCT/US24/23576 dated Sep. 6, 2024 (18 Pages). [cited by applicant]
Mittal et al. “Al-assisted Tagging of Deepfake Audio Calls using Challenge-Response.” Published in ARXIV on Feb. 28, 2024, pp. 1-18, Retrieved on Aug. 14, 2024, Retrieved from: https://arxiv.org/abs/2402.18085. [cited by applicant]
International Search Report and Written Opinion on International Application No. PCT/US 24/26212 dated Oct. 18, 2024 (19 Pages). [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/023576 dated Oct. 30, 2025. [cited by applicant]
International Preliminary Report on Patentability (including WO-ISA) from PCT/US2024/026212 dated Nov. 6, 2025. [cited by applicant]