IP Library Granted Patent US 12,417,772
Granted Patent B2
US 12,417,772 · App. 18/394,300 · Granted Sep 16, 2025

Robust spoofing detection system using deep residual neural networks

Inventors: Tianxiang Chen (Atlanta, GA); Elie Khoury (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L17/18G10L17/02G10L17/04G10L17/08G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,772
App. No.
18/394,300
Granted
Sep 16, 2025
Kind
B2
Abstract

Embodiments described herein provide for systems and methods for implementing a neural network architecture for spoof detection in audio signals. The neural network architecture contains a layers defining embedding extractors that extract embeddings from input audio signals. Spoofprint embeddings are generated for particular system enrollees to detect attempts to spoof the enrollee's voice. Optionally, voiceprint embeddings are generated for the system enrollees to recognize the enrollee's voice. The voiceprints are extracted using features related to the enrollee's voice. The spoofprints are extracted using features related to features of how the enrollee speaks and other artifacts. The spoofprints facilitate detection of efforts to fool voice biometrics using synthesized speech (e.g., deepfakes) that spoof and emulate the enrollee's voice.

Claims (37)

1. A computer-implemented method for spoofing countermeasures, the method comprising:

obtaining, by a computer, a plurality of training audio signals, including one or more clean training signals having one or more speaker biometric features and one or more spoof training signals having one or more spoofing features;

training, by the computer, a voiceprint embedding extractor for extracting a voiceprint based upon the one or more speaker biometric features, by applying the voiceprint embedding extractor on the one or more clean training signals having the one or more speaker biometric features;

training, by the computer, a spoofprint embedding extractor for extracting a spoofprint distinct from the voiceprint based upon the one or more spoofing features, by applying the spoofprint embedding extractor on the one or more spoof training signals having the one or more spoofing features; and

for an inbound audio signal at a deployment time, extracting, by the computer, an inbound spoofprint based upon the one or more spoofing features of the inbound audio signal by applying the spoofprint embedding extractor on the inbound audio signal.

2. The method according to claim 1 , wherein the computer trains the spoofprint embedding extractor using at least one clean audio signal.

3. The method according to claim 1 , wherein obtaining the plurality of training audio signals includes generating, by the computer, a spoof training signal as a simulated audio signal of a corresponding clean audio signal by executing one or more data augmentation operations on the corresponding clean audio signal.

4. The method according to claim 1 , wherein training the spoofprint embedding extractor includes:

applying, by the computer, the spoofprint embedding extractor on a spoof training signal to extract a training spoofprint based on the one or more spoofing features of the spoof training signal; and

executing, by the computer, a loss function using the training spoofprint extracted by the spoofprint embedding extractor for the spoof training signal.

5. The method according to claim 1 , further comprising updating, by the computer, one or more parameters of the spoofprint embedding extractor based upon a difference between the training spoofprint and an expected training spoofprint.

6. The method according to claim 1 , further comprising training, by the computer, a classifier of a neural network architecture for speaker recognition by applying the neural network architecture on each training voiceprint.

7. The method according to claim 1 , further comprising training, by the computer, a classifier of a neural network architecture for spoof detection by applying the neural network architecture on each training spoofprint.

8. The method according to claim 1 , further comprising:

obtaining, by the computer, a plurality of enrollment audio signals associated with an enrolled user, including one or more clean enrollment signals having the one or more speaker biometric features for the enrolled user and one or more enrollment spoof signals having the one or more spoofing features;

applying, by the computer, the spoofprint embedding extractor on the one or more enrollment spoof signals to extract an enrolled spoofprint for the enrolled user based upon the one or more spoofing features of the one or more enrollment spoof signals.

9. The method according to claim 8 , further comprising for the inbound audio signal at the deployment time, generating, by the computer, a spoof score for the inbound audio signal based upon a distance between the inbound spoofprint and the enrolled spoofprint.

10. The method according to claim 8 , further comprising:

applying, by the computer, the voiceprint embedding extractor on the one or more clean enrollment signals to extract an enrolled voiceprint for the enrolled user based upon the one or more speaker biometric features for the enrolled user of the one or more clean enrollment signals; and

for the inbound audio signal at the deployment time, generating, by the computer, a speaker recognition score for the inbound audio signal based upon a second distance between an inbound voiceprint and the enrolled voiceprint.

11. A computer-implemented method for spoofing countermeasures, the method comprising:

extracting, by a computer, an input voiceprint for an input audio signal based upon one or more speaker biometric features of the input audio signal by applying a voiceprint embedding extractor for extracting a voiceprint based upon the one or more speaker biometric features;

extracting, by the computer, an input spoofprint for the input audio signal based upon one or more spoofing features of the input audio signal by applying a spoofprint embedding extractor for extracting a spoofprint distinct from the voiceprint based upon the one or more spoofing features;

generating, by the computer, an input combined embedding based upon the input voiceprint and the input spoofprint; and

generating, by the computer, a combined similarity score between the input combined embedding and a second combined embedding.

12. The method according to claim 11 , further comprising identifying, by the computer, the input audio signal as being at least one of genuine or fraudulent based upon the combined similarity score and one or more corresponding preconfigured threshold scores.

13. The method according to claim 11 , further comprising:

extracting, by the computer, an enrolled voiceprint for an enrolled user based upon the one or more speaker biometric features in an enrollment audio signal;

extracting, by the computer, an enrolled spoofprint for the enrolled user based upon the one or more spoofing features in at least one enrollment audio signal; and

generating, by the computer, the second combined embedding as an combined enrollment embedding for the enrolled user based upon the enrolled voiceprint and the enrolled spoofprint.

14. The method according to claim 13 , further comprising generating, by the computer, one or more simulated enrollment spoofing signals by executing one or more data augmentation operations on the enrollment audio signal.

15. The method according to claim 11 , further comprising generating, by the computer, a spoofing score for the input audio signal based upon a distance between the input spoofprint and a second spoofprint used for the second combined embedding.

16. The method according to claim 11 , further comprising generating, by the computer, a voice similarity score for the input audio signal based upon a distance between the input voiceprint and a second voiceprint used for the second combined embedding.

17. The method according to claim 11 , further comprising executing, by the computer, a loss function using the combined similarity score as a distance loss between the input combined embedding as a predicted combined embedding and the second combined embedding as an expected combined embedding indicated by a training label associated with the input audio signal.

18. The method according to claim 17 , further comprising training, by the computer, the voiceprint embedding extractor by applying the voiceprint embedding extractor on a training audio signal having the one or more speaker biometric features, wherein the loss function updates one or more parameters of the voiceprint embedding extractor based upon the distance loss.

19. The method according to claim 17 , further comprising training, by the computer, the spoofprint embedding extractor by applying the spoofprint embedding extractor on at least one training audio signal having the one or more spoofing features, wherein the loss function updates one or more parameters of the spoofprint embedding extractor based upon the distance loss.

20. The method according to claim 17 , further comprising generating, by the computer, one or more simulated spoofed training signals by executing one or more data augmentation operations on the training audio signal.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: CHEN, TIANXIANG; KHOURY, ELIE
To: PINDROP SECURITY, INC.
Reel/Frame 067619/0859 →
Continuity (4)
Continuation 17155851 · Jan 22, 2021
Provisional Application 63068670 · Aug 21, 2020
Provisional Application 62966473 · Jan 27, 2020
Related Publication 20240153510A1 · May 9, 2024
References Cited (63)
US 5442696A · Lindberg · 1995 [cited by examiner]
US 5570412A · LeBlanc · 1996 [cited by examiner]
US 5724404A · Garcia · 1998 [cited by examiner]
US 5825871A · Mark · 1998 [cited by examiner]
US 6041116A · Meyers · 2000 [cited by examiner]
US 6134448A · Shoji · 2000 [cited by examiner]
US 6654459B1 · Bala · 2003 [cited by examiner]
US 6735457B1 · Link, II · 2004 [cited by examiner]
US 6765531B2 · Anderson · 2004 [cited by examiner]
US 7787598B2 · Agapi · 2010 [cited by examiner]
US 8050393B2 · Apple · 2011 [cited by examiner]
US 8223755B2 · Jennings · 2012 [cited by examiner]
US 8311218B2 · Mehmood · 2012 [cited by examiner]
US 8385888B2 · Labrador · 2013 [cited by examiner]
US 9060057B1 · Danis · 2015 [cited by examiner]
US 9078143B2 · Rodriguez · 2015 [cited by examiner]
US 9704478B1 · Vitaladevuni · 2017 [cited by examiner]
US 10257591B2 · Gaubitch · 2019 [cited by examiner]
US 11862177B2 · Chen · 2024 [cited by examiner]
US 20020181448A1 · Uskela · 2002 [cited by examiner]
US 20030012358A1 · Kurtz · 2003 [cited by examiner]
US 20110051905A1 · Maria Poels · 2011 [cited by examiner]
US 20110123008A1 · Sarnowski · 2011 [cited by examiner]
US 20150120027A1 · Cote · 2015 [cited by examiner]
US 20160293185A1 · Cote · 2016 [cited by examiner]
US 20170222960A1 · Agarwal · 2017 [cited by examiner]
US 20170302794A1 · Spievak · 2017 [cited by examiner]
US 20170359362A1 · Kashi · 2017 [cited by examiner]
US 20180197547A1 · Shi · 2018 [cited by examiner]
US 20180254046A1 · Khoury · 2018 [cited by examiner]
US 20190228778A1 · Lesso · 2019 [cited by examiner]
US 20190228779A1 · Lesso · 2019 [cited by examiner]
US 20190354808A1 · Park et al. · 2019 [cited by applicant]
US 20200322377A1 · Lakhdhar · 2020 [cited by examiner]
US 20210110813A1 · Khoury · 2021 [cited by examiner]
US 20210233541A1 · Chen · 2021 [cited by examiner]
US 20220121868A1 · Chen · 2022 [cited by examiner]
JP H05143094A · 1993 [cited by applicant]
JP 2018508799A · 2018 [cited by applicant]
WO WO2019145708A1 · 2019 [cited by applicant]
WO WO2020003533A1 · 2020 [cited by applicant]
Alzantot et al., “Deep Residual Neural Networks for Audio Spoofing Detection”, Department of Computer Science, UCLA, INTERSPEECH 2019, Sep. 15-19, 2019, Graz, Austria, pp. 1078-1082 (5 pages). [cited by applicant]
Alzantot, M. et al., “Deep Residual Neural Networks for Audio Spoofing Detection”, INTERSPEECH 2019, Graz, Austria, pp. 1078-1082, https://www.researchgate.net/profile/Ziqi-Wang-16/publication/334161923_Deep_Residual_Ne… [cited by applicant]
Canadian Examination Report dated Oct. 16, 2019, issued in corresponding Canadian Application No. 3,032,807, 3 pages. [cited by applicant]
Examination Report No. 1 for Australian app. 2021212621 dated Mar. 1, 2023 (4 pages). [cited by applicant]
Final Office Action on U.S. Appl. No. 17/155,851 dated May 25, 2023. [cited by applicant]
First Examiner's Requisition for CA App. 3,168,248 dated Aug. 21, 2023 (4 pages). [cited by applicant]
International Preliminary Report on Patentability for PCT Appl. Ser. No. PCT/US2021/014633, Aug. 11, 2022 (9 pages). [cited by applicant]
International Search Report and Written Opinion for PCT Appl. Ser. No. PCT/US2021/014633 dated Apr. 7, 2021 (11 pages). [cited by applicant]
International Search Report issued in International Application No. PCT/US2017/044849 dated Jan. 11, 2018 (8 pages). [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 17/155,851 dated Feb. 9, 2023 (21 pages). [cited by applicant]
Notice of Allowance on U.S. Appl. No. 17/155,851 dated Aug. 25, 2023. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 17/155,851 dated Sep. 27, 2023 (8 pages). [cited by applicant]
Schulzrinne et al., “RTP Payload for DTMF Digits, Telephone Tones, and Telephony Signals” Columbia University, Dec. 2006, (50 Pages)<https://tools.ielf.org/html/rfc4733.>. [cited by applicant]
Snyder et al., “X-Vectors: Robust DNN Embeddings for Speaker Recognition,” Center for Language and Speech Processing & Human Language Technology Center of Excellence, The Johns Hopkins University, 2018. [cited by applicant]
US Final Office Action on U.S. Appl. No. 15/155,851 dated Nov. 8, 2022 (21 pages). [cited by applicant]
US Non Final Office Action on U.S. Appl. No. 17/155,851 dated May 10, 2022 (18 pages). [cited by applicant]
Cai Weicheng et al: “The DKU Replay Detection System for the ASVspoof 2019 Challenge: On Data Augmentation, Feature Representation, Classification, and Fusion”, INTERSPEECH 2019, Jul. 5, 2019 (Jul. 5, 2019), pp. 1023-10… [cited by applicant]
Extended European Search Report on EPO App. 21747446.9 dated Jan. 23, 2024 (9 pages). [cited by applicant]
Li Jiakang et al: “Joint Decision of Anti-Spoofing and Automatic Speaker Verification by Multi-Task Learning With Contrastive Loss”, IEEE Access, IEEE, USA, vol. 8, Jan. 6, 2020 (Jan. 6, 2020), pp. 7907-7915, XP01176556… [cited by applicant]
Li Jiakang et al: “Multi-task learning of deep neural networks for joint automatic speaker verification and spoofing detection”, 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conferen… [cited by applicant]
EPO Examination Report for European Application No. 21747446.9 mailing date Jan. 17, 2025, 4 pages. [cited by applicant]
JP Office Action for Application No. 2022543650 mailing date Mar. 25, 2025, 12 pages. [cited by applicant]