IP Library Granted Patent US 12,609,132
Granted Patent B2
US 12,609,132 · App. 17/953,156 · Granted Apr 21, 2026

Voice modification detection using physical models of speech production

Inventors: David Looney (Atlanta, GA); Nikolay D. Gaubitch (Atlanta, GA)
Assignee: Pindrop Security, Inc.
G10L25/51G10L15/063G10L15/22G10L25/90H04M3/436
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,132
App. No.
17/953,156
Granted
Apr 21, 2026
Kind
B2
Abstract

A computer may train a single-class machine learning using normal speech recordings. The machine learning model or any other model may estimate the normal range of parameters of a physical speech production model based on the normal speech recordings. For example, the computer may use a source-filter model of speech production, where voiced speech is represented by a pulse train and unvoiced speech by a random noise and a combination of the pulse train and the random noise is passed through an auto-regressive filter that emulates the human vocal tract. The computer leverages the fact that intentional modification of human voice introduces errors to source-filter model or any other physical model of speech production. The computer may identify anomalies in the physical model to generate a voice modification score for an audio signal. The voice modification score may indicate a degree of abnormality of human voice in the audio signal.

Claims (35)

1 . A computer-implemented method comprising:

extracting, by a computer, a first set of speech parameters representing attributes of one or more normal speech signals, by applying a physical speech model on one or more training audio signals including the one or more normal speech signals;

training, by the computer, a machine-learning model to generate a modified-voice score, by applying the machine-learning model on the first set of speech parameters extracted for the one or more normal speech signals of the one or more training audio signals;

extracting, by the computer, a second set of the speech parameters representing attributes of an inbound speech signal, by applying the physical speech model on an inbound audio signal including the inbound speech signal; and

generating, by the computer, the modified-voice score for the inbound signal by applying the machine-learning model on the second set of the speech parameters and the first set of the speech parameters.

2 . The method according to claim 1 , further comprising determining, by the computer, whether an inbound phone call associated with the inbound audio signal is fraudulent based upon the voice modification score.

3 . The method according to claim 1 , wherein generating the modified-voice score includes comparing, by the computer, the first set of the speech parameters against the second set of the speech parameters to determine one or more similarity scores.

4 . The method according to claim 1 , wherein the machine learning model includes a single-class model trained for detecting one or more audio signal anomalies for the modified-voice score.

5 . The method according to claim 1 , wherein a speech parameter of the speech parameters includes at least one of: a pitch parameter, a formant parameter, a residual parameter, or linear predictive coding (LPC) model order parameter.

6 . The method according to claim 5 , further comprising:

determining, by the computer, a first anomaly score based upon comparing the pitch parameter of the first set of the parameters for the attributes of the normal human speech signals against the pitch parameter of the second set of the parameters for the attributes of the inbound speech signal; and

determining, by the computer, a second anomaly score based upon comparing the formant parameter of the first set of the parameters for the attributes of the normal human speech signals against the formant parameter of the second set of the parameters for the attributes of the inbound speech signal,

wherein the computer determines the modified-voice score based upon each anomaly score.

7 . The method according to claim 1 , wherein the physical speech model applies a source-filter model including a pulse train and white random noise for extracting the speech parameters.

8 . The method according to claim 7 , wherein the source-filter model applies an LPC filter for extracting one or more formant parameters of the speech parameters.

9 . The method according to claim 7 , wherein the source-filter model applies an inverse LPC filter for extracting one or more residual parameters of the speech parameters.

10 . The method according to claim 9 , further comprising generating, by the computer, an LPC filter based upon the one or more residual parameters generated using the inverse LPC filter.

11 . A system comprising:

a computer comprising a processor configured to:

extract a first set of speech parameters representing attributes of one or more normal speech signals, by applying a physical speech model on one or more training audio signals including the one or more normal speech signals;

train a machine-learning model to generate a modified-voice score, by applying the machine-learning model on the first set of speech parameters extracted for the one or more normal speech signals of the one or more training audio signals;

extract a second set of the speech parameters representing attributes of an inbound speech signal, by applying the physical speech model on an inbound audio signal including the inbound speech signal; and

generate the modified-voice score for the inbound signal by applying the machine-learning model on the second set of the speech parameters and the first set of the speech parameters.

12 . The system according to claim 11 , wherein the computer is further configured to determine whether an inbound phone call associated with the inbound audio signal is fraudulent based upon the voice modification score.

13 . The system according to claim 11 , wherein, when generating the modified-voice score, the computer is further configured to compare the first set of the speech parameters against the second set of the speech parameters to determine one or more anomaly scores.

14 . The system according to claim 11 , wherein the machine learning model includes a single-class model trained for detecting one or more audio signal anomalies for the modified-voice score.

15 . The system according to claim 11 , wherein a speech parameter of the speech parameters includes at least one of: a pitch parameter, a formant parameter, a residual parameter, or linear predictive coding (LPC) model order parameter.

16 . The system according to claim 15 , wherein the computer is further configured to:

determine a first anomaly score based upon comparing the pitch parameter of the first set of the parameters for the attributes of the normal human speech signals against the pitch parameter of the second set of the parameters for the attributes of the inbound speech signal; and

determine a second anomaly score based upon comparing the formant parameter of the first set of the parameters for the attributes of the normal human speech signals against the formant parameter of the second set of the parameters for the attributes of the inbound speech signal,

wherein the computer determines the modified-voice score based upon each anomaly score.

17 . The system according to claim 11 , wherein the physical speech model applies a source-filter model including a pulse train and white random noise for extracting the speech parameters.

18 . The system according to claim 17 , wherein the source-filter model applies an LPC filter for extracting one or more formant parameters of the speech parameters.

19 . The system according to claim 17 , wherein the source-filter model applies an inverse LPC filter for extracting one or more residual parameters of the speech parameters.

20 . The system according to claim 19 , wherein the computer is further configured to generate an LPC filter based upon the one or more residual parameters generated using the inverse LPC filter.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
RELEASE OF SECURITY INTEREST Recorded Jun 26, 2024
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: PINDROP SECURITY, INC.
Reel/Frame 069477/0962 →
SECURITY INTEREST Recorded Jul 31, 2023
From: PINDROP SECURITY, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064443/0584 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2022
From: LOONEY, DAVID; GAUBITCH, NIKOLAY D.
To: PINDROP SECURITY, INC.
Reel/Frame 061230/0745 →
Continuity (3)
Continuation 16375785 · Apr 4, 2019
Provisional Application 62652539 · Apr 4, 2018
Related Publication 20230015189A1 · Jan 19, 2023
References Cited (33)
US 6006188A · Bogdashevsky et al. · 1999 [cited by applicant]
US 6256609B1 · Byrnes · 2001 [cited by examiner]
US 8682650B2 · Gray et al. · 2014 [cited by applicant]
US 9372976B2 · Bukai · 2016 [cited by applicant]
US 9786275B2 · Coifman et al. · 2017 [cited by applicant]
US 9786300B2 · Chan et al. · 2017 [cited by applicant]
US 11495244B2 · Looney · 2022 [cited by examiner]
US 11646018B2 · Maddali · 2023 [cited by examiner]
US 20010056349A1 · St. John · 2001 [cited by examiner]
US 20020002460A1 · Pertrushin · 2002 [cited by applicant]
US 20020010587A1 · Pertrushin · 2002 [cited by examiner]
US 20100228656A1 · Wasserblat et al. · 2010 [cited by applicant]
US 20130226569A1 · Chandra et al. · 2013 [cited by applicant]
US 20140214676A1 · Bukai · 2014 [cited by applicant]
US 20150142446A1 · Gopinathan et al. · 2015 [cited by applicant]
US 20180146370A1 · Krishnaswamy · 2018 [cited by examiner]
US 20200312313A1 · Maddali · 2020 [cited by examiner]
WO WO2011083362A1 · 2011 [cited by applicant]
Chan et al., “Modeling Multiple Time Series for Anomaly Detection”, Retrieved from the Internet: https://www.researchgate.net/publication/220766749, Conference Paper—Jan. 2005. [cited by applicant]
International Preliminary Report On Patentability and Written Opinion of the International Searching Authority issued in PCT/US2019/025893 with Date of mailing Jul. 9, 2019. [cited by applicant]
International Search Report and the Written Opinion of the International Searching Authority dated Jul. 9, 2019, issued in corresponding International Application No. PCT/US2019/025893, 12 pages. [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 16/375,785 DTD Feb. 10, 2022. [cited by applicant]
Notice of Allowance on U.S. Appl. No. 16/375,785 DTD Jul. 1, 2022. [cited by applicant]
Wu et al., “Synthetic Speech Detection Using Temporal Modulcation Feature”, http://www.zhizheng.org/papers/CASS2013_synthetic_dection.pdf.Publication date via https://ieeexplore.ieee.org/document/6639067, Date of Confer… [cited by applicant]
Adiga, et al., “Improved voicing decision using glottal activity features for statistical parametric speech synthesis.” Digital Signal Processing, 71 (2017), pp. 131-143. [cited by applicant]
Alegre, F., et al., “A one-class classification approach to generalized speaker verification spoofing countermeasures using local binary patterns”, in 2013 IEEE Sixth International Conference on Biometrics: Theory, Appl… [cited by applicant]
Aurino, F., et al., “One-class SVM based approach for detecting anomalous audio events,” In 2014 International Conference on Intelligent Networking and Collaborative Systems, pp. 145-151, IEEE, 2014. [cited by applicant]
E. Nemer, et al., “Robust voice activity detection using higher-order statistics in the LPC residual domain,” in IEEE Transactions on Speech and Audio Processing, vol. 9, No. 3, pp. 217-231, Mar. 2001. [cited by applicant]
S. Safavi, et al., “Fraud Detection in Voice-Based Identity Authentication Applications and Services,” 2016 IEEE 16th International Conference on Data Mining Workshops, pp. 1074-1081. [cited by applicant]
Srivastava, S. et al., “Formant based linear prediction coefficients for speaker identification,” In 2014 International Conference on Signal Processing and Integrated Networks (SPIN) (pp. 685-688), IEEE, 2014. [cited by applicant]
V.S. Baidwan, et al., “Comparative Analysis of Prosodic Features and Linear Predictive Coefficients for Speaker Recognition Using Machine Learning Technique,” 2014 International Conference on Devices, Circuits and Commu… [cited by applicant]
Vaquero, C., et al., “On the need of template protection for voice authentication. In Sixteenth Annual Conference of the International Speech Communication Association”, 2015. [cited by applicant]
Zhou, Y., et al., “Speaker verification based on SVDD. In 2010, 3rd International Congress on Image and Signal Processing,” vol. 7, pp. 3168-3172), IEEE, 2010. [cited by applicant]