IP Library Granted Patent US 12,567,401
Granted Patent B2
US 12,567,401 · App. 18/301,064 · Granted Mar 3, 2026

Evaluating reliability of audio data for use in speech processing

Inventors: Sarah Bakst (San Francisco, CA); Aaron Lawson (Sutter Creek, CA); Christopher L. Cobo-Kroenke (San Francisco, CA); Allen Stauffer (Queensbury, NY)
Assignee: SRI INTERNATIONAL
G10L15/02G10L15/063G10L15/08G10L15/28G10L2015/081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,401
App. No.
18/301,064
Granted
Mar 3, 2026
Kind
B2
Abstract

In some examples, a computing system includes a storage device configured to store a machine learning model trained with audio feature values to determine a reliability of an audio segment for performing speech processing; and processing circuitry. The processing circuitry is configured to: receive an audio dataset comprising a sequence of audio segments; extract, for each audio segment of the sequence of audio segments, a set of audio feature values corresponding to a set of audio features; execute the machine learning model to determine, for each audio segment of the sequence of audio segments, a reliability score based on the set of audio feature values corresponding to the respective audio segment, wherein the reliability score indicates a reliability of the audio segment for performing speech processing; and output an indication of the respective reliability scores determined for at least one audio segment of the sequence of audio segments.

Claims (81)

1 . A computing system comprising:

a storage device configured to store a machine learning model trained with audio feature values to determine a reliability of an audio segment for performing speech processing; and

processing circuitry having access to the storage device and configured to:

receive an audio dataset comprising a sequence of audio segments including speech from an unknown speaker;

extract, for each audio segment of the sequence of audio segments, a set of audio feature values corresponding to a set of audio features, wherein each audio feature value of the set of audio feature values indicates a prevalence of the corresponding audio feature of the set of audio features in the audio segment;

execute the machine learning model to determine, for each audio segment of the sequence of audio segments, a reliability score based on the set of audio feature values corresponding to the respective audio segment, wherein the reliability score indicates a reliability of the audio segment for performing speech processing;

select, based on the reliability score for each audio segment, one or more audio segments from the sequence of audio segments;

store the one or more audio segments as most valuable for performing speaker identification;

determine, based on the one or more audio segments, a probability indicating whether the speech from the unknown speaker corresponds to a target speaker; and

output an indication of the probability indicating whether the speech from the unknown speaker corresponds to the target speaker.

2 . The computing system of claim 1 , wherein to select the one or more audio segments from the sequence of audio segments, the processing circuitry is configured to:

identify, based on the respective reliability scores for the sequence of audio segments, one or more audio segments of the sequence of audio segments that are most valuable for performing speech processing.

3 . The computing system of claim 1 ,

wherein the storage device is further configured to store a speech processing model,

wherein the speech from the unknown speaker included in the audio dataset is associated with an unknown target class, and

wherein to determine the probability, the processing circuitry is further configured to:

determine, based on the one or more audio segments of the sequence of audio segments that are most valuable for performing speech processing and using the speech processing model, whether the unknown target class corresponding to the audio dataset is the same as a known target class corresponding to one or more reference audio datasets associated with the target speaker; and

determine the probability based on information indicating whether the unknown target class corresponding to the audio dataset is the same as the known target class.

4 . The computing system of claim 1 , wherein the set of audio features includes any one or more of cycle-to-cycle changes in amplitude, cycle-to-cycle changes in frequency, signal-to-noise ratio (SNR), harmonics-to-noise ratio (HNR), degradation due to reverberation, emotional valence, level of speech activity, voicing probability, autocorrelation peak, mean spectral tilt, and standard deviation of spectral tilt.

5 . The computing system of claim 1 , wherein to extract the set of audio feature values for each audio segment of the sequence of audio segments, the processing circuitry is configured to:

calculate, based on a portion of the audio dataset corresponding to each audio segment of the sequence of audio segments, an audio feature value corresponding to each audio feature of the set of the set of audio features.

6 . The computing system of claim 5 , wherein the machine learning model is a first machine learning model, and wherein to calculate an audio feature value corresponding to each audio feature of the set of the set of audio features, the processing circuitry is configured to:

calculate, using a feature extraction signal processing model, a first one or more audio feature values corresponding to the respective audio segment; and

execute one or more second machine learning models to determine, based on respective audio segment, a second one or more audio feature values.

7 . The computing system of claim 1 , wherein the audio dataset extends for a first duration of time, and wherein each audio segment of the sequence of audio segments comprises a portion of the audio dataset extending for a second duration of time that is shorter than the first duration of time.

8 . The computing system of claim 1 , wherein the processing circuitry is further configured to train the machine learning model.

9 . The computing system of claim 8 , wherein the storage device is further configured to store a speech processing model, wherein the storage device is configured to store training data comprising a plurality of training audio datasets, wherein each training audio dataset of the plurality of training audio datasets is known to include speech from a particular target class, wherein each training audio dataset of the plurality of training audio datasets comprises a sequence of training audio segments, and wherein to train the machine learning model, the processing circuitry is configured to:

execute a speech processing model to perform a plurality of speech processing determinations, wherein each speech processing determinations of the plurality of speech processing determinations involves determining whether a training audio segment of a training audio dataset corresponds to a particular target class;

determine, based on the particular target class known to be associated with each training audio dataset of the plurality of training audio datasets, a quality of each speech processing determination of the plurality of speech processing determinations;

identify one or more audio features of a plurality of audio features present in the training audio segment corresponding to each speech processing determination of the plurality of speech processing determinations; and

select the set of audio features from the plurality of audio features based on the quality of each speech processing determination of the plurality of speech processing determinations and the one or more audio features of the plurality of audio features present in the audio segment corresponding to each speech processing determination of the plurality of speech processing determinations.

10 . A method comprising:

receiving, by processing circuitry having access to a storage device configured to store a machine learning model trained with audio feature values to determine a reliability of an audio segment for performing speech processing, an audio dataset comprising a sequence of audio segments including speech from an unknown speaker;

extracting, by the processing circuitry for each audio segment of the sequence of audio segments, a set of audio feature values corresponding to a set of audio features, wherein each audio feature value of the set of audio feature values indicates a prevalence of the corresponding audio feature of the set of audio features in the audio segment;

executing, by the processing circuitry, the machine learning model to determine, for each audio segment of the sequence of audio segments, a reliability score based on the set of audio feature values corresponding to the respective audio segment, wherein the reliability score indicates a reliability of the audio segment for performing speech processing;

selecting, by the processing circuitry, based on the reliability score for each audio segment, one or more audio segments from the sequence of audio segments;

storing, by the processing circuitry, the one or more audio segments as most valuable for performing speaker identification;

determining, by the processing circuitry, based on the one or more audio segments, a probability indicating whether the speech from the unknown speaker corresponds to a target speaker; and

outputting, by the processing circuitry, an indication of the probability indicating whether the speech from the unknown speaker corresponds to the target speaker.

11 . The method of claim 10 , wherein selecting the one or more audio segments from the sequence of audio segments comprises:

identifying, by the processing circuitry based on the respective reliability scores for the sequence of audio segments, one or more audio segments of the sequence of audio segments that are most valuable for performing speech processing.

12 . The method of claim 10 , wherein the storage device is further configured to store a speech processing model, wherein the speech from the unknown speaker included in the audio dataset is associated with an unknown target class, and wherein determining the probability comprises:

determining, by the processing circuitry, based on the one or more audio segments of the sequence of audio segments that are most valuable for performing speech processing and using the speech processing model, whether the unknown target class corresponding to the audio dataset is the same as a known target class corresponding to one or more reference audio datasets associated with the target speaker; and

determining, by the processing circuitry, the probability based on information indicating whether the unknown target class corresponding to the audio dataset is the same as the known target class.

13 . The method of claim 10 , wherein extracting the set of audio feature values for each audio segment of the sequence of audio segments of the audio dataset comprises:

calculating, based on a portion of the audio dataset corresponding to each audio segment of the sequence of audio segments, an audio feature value corresponding to each audio feature of the set of the set of audio features.

14 . The method of claim 13 , wherein the machine learning model is a first machine learning model, and wherein calculating an audio feature value corresponding to each audio feature of the set of the set of audio features comprises:

calculating, by the processing circuitry, using a feature extraction signal processing model, a first one or more audio feature values corresponding to the respective audio segment; and

executing, by the processing circuitry, one or more second machine learning models to determine, based on respective audio segment, a second one or more audio feature values.

15 . The method of claim 10 , wherein the audio dataset extends for a first duration of time, and wherein each audio segment of the sequence of audio segments comprises a portion of the audio dataset extending for a second duration of time that is shorter than the first duration of time.

16 . The method of claim 15 , wherein the second duration of time corresponding to each audio segment of the sequence of audio segments comprises two seconds, and wherein the sequence of audio segments comprises a sequence of non-overlapping two second time windows extending for a length of the audio dataset.

17 . The method of claim 10 , further comprising training, by the processing circuitry, the machine learning model.

18 . The method of claim 17 , wherein the storage device is further configured to store a speech processing model, wherein the storage device is configured to store training data comprising a plurality of training audio datasets, wherein each training audio dataset of the plurality of training audio datasets is known to include speech from a particular target class, wherein each training audio dataset of the plurality of training audio datasets comprises a sequence of training audio segments, and wherein training the machine learning model comprises:

executing, by the processing circuitry, a speech processing model to perform a plurality of speech processing determinations, wherein each speech processing determination of the plurality of speech processing determinations involves determining whether a training audio segment of a training audio dataset corresponds to a particular target class;

determining, by the processing circuitry based on the particular target class known to be associated with each training audio dataset of the plurality of training audio datasets, a quality of each speech processing determination of the plurality of speech processing determinations;

identifying, by the processing circuitry, one or more audio features of a plurality of audio features present in the training audio segment corresponding to each speech processing determination of the plurality of speech processing determinations; and

selecting, by the processing circuitry, the set of audio features from the plurality of audio features based on the quality of each speech processing determination of the plurality of speech processing determinations and the one or more audio features of the plurality of audio features present in the audio segment corresponding to each speech processing determination of the plurality of speech processing determinations.

19 . Non-transitory computer-readable media medium comprising instructions that, when executed by a processor, cause the processor to:

receive an audio dataset comprising a sequence of audio segments including speech from an unknown speaker;

extract, for each audio segment of the sequence of audio segments, a set of audio feature values corresponding to a set of audio features, wherein each audio feature value of the set of audio feature values indicates a prevalence of the corresponding audio feature of the set of audio features in the audio segment;

execute a machine learning model to determine, for each audio segment of the sequence of audio segments, a reliability score based on the set of audio feature values corresponding to the respective audio segment, wherein the reliability score indicates a reliability of the audio segment for performing speech processing;

select, based on the reliability score for each audio segment, one or more audio segments from the sequence of audio segments;

store the one or more audio segments as most valuable for performing speaker identification;

determine, based on the one or more audio segments, a probability indicating whether the speech from the unknown speaker corresponds to a target speaker; and

output an indication of the probability indicating whether the speech from the unknown speaker corresponds to the target speaker.

20 . A computing system comprising:

a storage device configured to store a machine learning model and training data comprising a plurality of training audio datasets, wherein each training audio dataset of the plurality of training audio datasets is known to include speech from a particular target class, wherein each training audio dataset of the plurality of training audio datasets comprises a sequence of training audio segments; and

processing circuitry having access to the storage device, configured to determine a probability indicating whether speech from an unknown speaker corresponds to a target speaker, and configured to train the machine learning model to determine a reliability of an audio segment for performing speech processing,

wherein to train the machine learning model, the processing circuitry is configured to:

execute a speech processing model to perform a plurality of speech processing determinations, wherein each speech processing determination of the plurality of speech processing determinations involves determining whether a training audio segment of a training audio dataset corresponds to a particular target class;

determine, based on the particular target class known to be associated with each training audio dataset of the plurality of training audio datasets, a quality of each speech processing determination of the plurality of speech processing determinations;

identify one or more audio features of a plurality of audio features present in the training audio segment corresponding to each speech processing determination of the plurality of speech processing determinations; and

select a set of audio features of the plurality of audio features based on the quality of each speech processing determination of the plurality of speech processing determinations and the one or more audio features of the plurality of audio features present in the audio segment corresponding to each speech processing determination of the plurality of speech processing determinations, and

wherein to determine the probability indicating whether speech from the unknown speaker corresponds to the target speaker, the processing circuitry is configured to:

receive an audio dataset comprising a sequence of audio segments including speech from the unknown speaker;

extract, for each audio segment of the sequence of audio segments, a set of audio feature values corresponding to the set of audio features, wherein each audio feature value of the set of audio feature values indicates a prevalence of the corresponding audio feature of the set of audio features in the audio segment;

execute the machine learning model to determine, for each audio segment of the sequence of audio segments, a reliability score based on the set of audio feature values corresponding to the respective audio segment, wherein the reliability score indicates a reliability of the audio segment for performing speech processing;

select, based on the reliability score for each audio segment, one or more audio segments from the sequence of audio segments;

store the one or more audio segments as most valuable for performing speaker identification;

determine, based on the one or more audio segments, a probability indicating whether the speech from the unknown speaker corresponds to the target speaker; and

output an indication of the probability indicating whether the speech from the unknown speaker corresponds to the target speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2023
From: BAKST, SARAH; LAWSON, AARON; COBO-KROENKE, CHRISTOPHER L.; STAUFFER, ALLEN
To: SRI INTERNATIONAL
Reel/Frame 064838/0566 →
Continuity (2)
Provisional Application 63331713 · Apr 15, 2022
Related Publication 20230335114A1 · Oct 19, 2023
References Cited (97)
US 9147401B2 · Shriberg et al. · 2015 [cited by applicant]
US 10133538B2 · McLaren et al. · 2018 [cited by applicant]
US 10476872B2 · McLaren · 2019 [cited by examiner]
US 11024291B2 · Castan Lavilla et al. · 2021 [cited by applicant]
US 11823658B2 · McLaren · 2023 [cited by examiner]
US 12035106B2 · Lawson · 2024 [cited by examiner]
US 20140278412A1 · Scheffer et al. · 2014 [cited by applicant]
US 20160248768A1 · Mclaren et al. · 2016 [cited by applicant]
US 20160283185A1 · Mclaren et al. · 2016 [cited by applicant]
US 20190013013A1 · McLaren et al. · 2019 [cited by applicant]
US 20190180771A1 · Yin · 2019 [cited by examiner]
US 20210089907A1 · Rogers et al. · 2021 [cited by applicant]
US 20210125629A1 · Bryan · 2021 [cited by examiner]
US 20220044077A1 · Lawson et al. · 2022 [cited by applicant]
US 20220093106A1 · Mosayyebpour Kaskari · 2022 [cited by examiner]
US 20220111294A1 · Barreiro · 2022 [cited by examiner]
US 20230053148A1 · Yu · 2023 [cited by examiner]
US 20230401285A1 · Kobren · 2023 [cited by examiner]
US 20250316062A1 · Salamon · 2025 [cited by examiner]
Boles et al., “Voice Biometrics: Deep Learning-based Voiceprint Authentication System” IEEE Conferences 2017 12th System of System Engineering Conference) SoSE01/06/2017 (Year: 2017). [cited by examiner]
Beck et al., “A Bilingual Multi-Modal Voice Corpus for Language and Speaker Recognition (LASR) Services”, The Speaker and Language Recognition Workshop, May 31, 2004, 6 pp. [cited by applicant]
Boril et al., “Unsupervised Equalization of Lombard Effect for Speech Recognition in Noisy Adverse Environments,” IEEE Transactions Audio, Speech, and Language Processing, vol. 18, No. 6, Aug. 2010, pp. 1379-1393. [cited by applicant]
Bou-Ghazale et al., “A Comparative Study of Traditional and Newly Proposed Features for Recognition of Speech Under Stress,” IEEE Transactions on Speech & Audio Processing, vol. 8, No. 4, Jul. 2000, pp. 429-442. [cited by applicant]
Bou-Ghazale et al., “HMM-Based Stressed Speech Modeling with Application to Improved Synthesis and Recognition of Isolated Speech Under Stress,” IEEE Transactions on Speech & Audio Processing, vol. 6, No. 3, May 1998, p… [cited by applicant]
Campbell et al., “Estimating and evaluating confidence for forensic speaker recognition”, Proceedings. (ICASSP '05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005, Mar. 23, 2005, Page?. [cited by applicant]
Espy-Wilson et al., “A New Set of Features for Text-Independent Speaker Identification”, Ninth International Conference on Spoken Language Processing, Sep. 17, 2006, 4 pp. [cited by applicant]
Fan et al., “Acoustic Analysis and Feature Transformation from Neutral to Whisper for Speaker Identification within Whispered Speech Audio Streams,” Speech Communication, vol. 55, Jan. 2013, pp. 119-134. [cited by applicant]
Fan et al., “Speaker Identification within Whispered Speech Audio Streams,” IEEE Transactions Audio, Speech and Language Processing, vol. 19, No. 5, Jul. 2011, pp. 1408-1421. [cited by applicant]
Fenu et al., “Improving Fairness in Speaker Recognition”, Proceedings of the 2020 European Symposium on Software Engineering, Nov. 2020, 8 pp. [cited by applicant]
Ferrer et al., “A Noise-Robust System for NIST 2012 Speaker Recognition Evaluation,” SRI International Menlo Park CA Speech Technology and Research Laboratory, Aug. 2013, 6 pp. [cited by applicant]
Ferrer et al., “A Speaker Verification Backend with Robust Performance across Conditions”, Computer Science and Language, vol. 71, Aug. 18, 2021, 53 pp. [cited by applicant]
Ferrer et al., “Classification of Lexical Stress using Spectral and Prosodic Features for Computer-Assisted Language Learning Systems”, Speech Communication, vol. 69, May 2015, 20 pp. [cited by applicant]
Ferrer et al., “A Unified Approach for Audio Characterization and its Application to Speaker Recognition,” Odyssey 2012—The Speaker and Language Recognition Workshop, Singapore, 2012, 7 pp. (Applicant points out, in acc… [cited by applicant]
Ferrer et al., “Promoting robustness for speaker modeling in the community: the PRISM evaluation set,” Proceedings of NIST 2011 workshop, Atlanta, Dec. 2011, 7 pp. [cited by applicant]
Garcia-Romero et al., “Analysis of I-vector Length Normalization in Speaker Recognition Systems”, Twelfth annual conference of the international speech communication association, Aug. 2011, pp. 249-252. [cited by applicant]
Ghaffarzadegan et al., “Generative Modeling of Pseudo-Whisper for Robust Whispered Speech Recognition,” IEEE Transactions Audio, Speech, and Language Processing, vol. 24, No. 10, Oct. 2016, pp. 1705-1720. [cited by applicant]
Ghaffarzadegan et al., “UT-Vocal Effort II: Analysis and Constrained-Lexicon Recognition of Whispered Speech,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy, May… [cited by applicant]
Gillespie et al., “Speech Dereverberation via Maximum-Kurtosis Subband Adaptive Filtering”, 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 01CH37221), vol. 6, May 7… [cited by applicant]
Godin et al., “Analysis of the effects of physical task stress on the speech signal,” Journal of the Acoustical Society of America, vol. 130, No. 6, Dec. 2011, pp. 3992-3998. [cited by applicant]
Hansen et al., “Analysis and Compensation of Lombard Speech Across Noise Type and Levels With Application to In-Set/Out-of-Set Speaker Recognition,” IEEE Transactions Audio, Speech & Language Processing, vol. 17, No. 2,… [cited by applicant]
Hansen et al., “Feature Analysis and Neural Network based Classification of Speech under Stress,” IEEE Transactions on Speech & Audio Processing, vol. 4, No. 4, Jul. 1996, pp. 307-313. [cited by applicant]
Hansen et al., “Speaker Recognition by Machines and Humans: A Tutorial Review,” IEEE Signal Processing Magazine, Nov. 2015, pp. 74-99. [cited by applicant]
Hansen et al., “TEO-based Speaker Stress Assessment using Hybrid Classification and Tracking Schemes,” International Journal Speech Technology, vol. 15, Issue 3, Sep. 2012, pp. 295-311. [cited by applicant]
Hansen, “Analysis and Compensation of Speech under Stress and Noise for Environmental Robustness in Speech Recognition,” Speech Communication, Special Issue on Speech Under Stress, vol. 20(2), Nov. 1996, pp. 151-173. [cited by applicant]
Hansen, “Getting Started with SUSAS: A Speech Under Simulated and Actual Stress Database,” Eurospeech, vol. 4, Rhodes, Greece, Sep. 1997, 4 pp. [cited by applicant]
Hansen, “Robust Emotional Stressed Speech Detection using Weighted Frequency Subbands,” EURASIP Journal on Advances in Signal Processing: Special Issue on Emotion and Mental State Recognition from Speech, Apr. 2011, 10 … [cited by applicant]
Hansen, et al., “Driver Modeling for Detection & Assessment of Driver Distraction: Examples from the UTDrive Test Bed,” IEEE Signal Processing Magazine, Jul. 2017, 13 pp. [cited by applicant]
Hayashida et al., “Close/distant talker discrimination based on kurtosis of linear prediction residual signals”, 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 4, 2014, 5 pp. [cited by applicant]
He et al., “Between-speaker variability and temporal organization of the first formant”, The Journal of the Acoustical Society of America, vol. 145, No. 3, Mar. 11, 2019, EL209-EL214 pp. [cited by applicant]
Hicklin et al., “Assessing the clarity of friction ridge impressions”, Forensic Science International, vol. 226, No. 1-3, Mar. 10, 2013, 106-117 pp. [cited by applicant]
Huggins et al., “Confidence Metrics for Speaker Identification”, Seventh International Conference on Spoken Language Processing, Sep. 16, 2002, 4 pp. [cited by applicant]
Kaushik et al., “Automatic Sentiment Detection in Naturalistic Audio,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 25, No. 8, Aug. 2017, pp. 1668-1679. [cited by applicant]
Kaushik et al., “Automatic Audio Sentiment Extraction Using Keyword Spotting,” Sixteenth Annual Conference of the International Speech Communication Association, Sep. 6-10, 2015, 5 pp. [cited by applicant]
Lawson et al., “Long Term Examination of Intra-Session and Inter-Session Speaker Variability,” Research Associates for Defense Conversion (RADC), Marcy, NY, Mar. 2009, 4 pp. [cited by applicant]
Lei et al., “A Deep Neural Network Speaker Verification System Targeting Microphone Speech,” Fifteenth Annual Conference of the International Speech Communication Association, Sep. 2014, pp. 681-685. [cited by applicant]
Lei et al., “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2014, 6 pp. [cited by applicant]
Lei et al., “Simplified VTS-based i-vector extraction in noise-robust speaker recognition,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2014. 5 pp. [cited by applicant]
Lu et al., “The Effect of Language Factors for Robust Speaker Recognition”, 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19, 2009, 4217-4220 pp. [cited by applicant]
Mariooryad et al., “Exploring cross-modality affective reactions for audiovisual emotion recognition,” IEEE Transactions on Affective Computing, vol. 4, No. 2, Jan. 2013, 15 pp. [cited by applicant]
Mcdougall et al., “Discrimination of Speakers Using the Formant Dynamics of /u:/ in British English”, Proceedings of the International Congress of Phonetic Sciences, Aug. 6, 2007, 1825-1828 pp. [cited by applicant]
McLaren et al. “Combining Continuous Progressive Model adaptation and Factor Analysis for Speaker Verification,” Proceedings of the 9th Annual Conference of the International Speech Communication Association (Interspeec… [cited by applicant]
McLaren et al. “Application of convolutional neural networks to speaker recognition in noisy conditions,” Fifteenth Annual Conference of the International Speech Communication Association, Sep. 2014, pp. 686-690. [cited by applicant]
McLaren et al., “A Comparison of Session Variability Compensation Approaches for Speaker Verification” IEEE Transactions on Information Forensics and Security, vol. 5. No. 4, Aug. 2010, 8 pp. [cited by applicant]
McLaren et al., “Advances in Deep Neural Network Approaches to Speaker Recognition,” 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), Apr. 2015, pp. 4814-4818. [cited by applicant]
McLaren et al., “Exploring the Role of Phonetic Bottleneck Features for Speaker and Language Recognition,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5575-5579. [cited by applicant]
Mclaren et al., “How to Train Your Speaker Embeddings Extractor”, The Speaker and Language Recognition Workshop (Odyssey 2018), Jun. 26, 2018, pp. 327-334. [cited by applicant]
McLaren et al., “Improved Speaker Recognition Using DCT Coefficients as Features,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2015, pp. 4430-4434. [cited by applicant]
McLaren et al., “On the Issue of Calibration in DNNbased Speaker Recognition Systems,” Interspeech, Sep. 2016, pp. 1825-1829. [cited by applicant]
Mclaren et al., “Softsad: Integrated Frame-Based Speech Confidence for Speaker Recognition”, IEEE International Conference on Acoustics, Apr. 19, 2015, 5 pp. [cited by applicant]
Mclaren et al., “Trial-Based Calibration for Speaker Recognition in Unseen Conditions”, Proc. Odyssey, Jun. 16, 2014, 19-25 pp. [cited by applicant]
McLaren et al., “Effective use of DCTs for Contextualizing Features for Speaker Recognition,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2014, pp. 4027-4031. [cited by applicant]
McLaren et al., “Improving robustness to compressed speech in speaker recognition,” Interspeech, Aug. 2013, pp. 3698-3702. [cited by applicant]
McLaren, “Data-driven Impostor Selection for T-norm score Normalisation and the Background Dataset in SVM-based Speaker Verification,” Advances in Biometrics, Jun. 2009, pp. 474-483. [cited by applicant]
Morrison et al., “An empirical estimate of the precision of likelihood ratios from a forensic-voice-comparison system,” Forensic Science International, vol. 208, May 2011, pp. 59-65. [cited by applicant]
Morrison et al., “Assessing the admissibility of a new generation of forensic voice comparison testimony,” Columbia Science and Technology Law Review, vol. 18, Spring 2016, pp. 326-434. [cited by applicant]
Park et al., “Using Voice Quality Features to Improve Short-Utterance, Text-Independent Speaker Verification Systems”, Interspeech, Aug. 20, 2017, 5 pp. [cited by applicant]
Parthasarathy et al., “Predicting Speaker Recognition Reliability by Considering Emotional Content”, 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII), Oct. 23, 2017, 434-43… [cited by applicant]
Parthasarathy et al., “Defining Emotionally Salient Regions using Qualitative Agreement Method,” Interspeech, San Francisco, CA, USA, Sep. 2016, pp. 3598-3602. [cited by applicant]
Patil et al., “The physiological microphone (PMIC): A competitive alternative for speaker assessment in stress detection and speaker verification,” Speech Communication: Special Issue on Silent Speech Interfaces, vol. 5… [cited by applicant]
Poh et al., “Estimating the confidence interval of expected performance curve in biometric authentication using joint bootstrap,” 2007 IEEE International Conference on Acoustics, Speech and Signal Processing—ICASSP'07, … [cited by applicant]
President's Council of Advisors on Science and Technology 2016 Report, “Forensic science in criminal courts: Ensuring scientific validity of feature-comparison methods,” Sep. 2016, 174 pp. [cited by applicant]
Prince et al., “Probabilistic Linear Discriminant Analysis for Inferences About Identity”, 2007 IEEE 11th International Conference on Computer Vision, Oct. 14, 2007, 8 pp. [cited by applicant]
R Core Team, “The R Project for Statistical Computing”, R Foundation for Statistical Computing, Vienna, Austria, Retrieved from: https://www.r-project.org/, Accessed on : Jul. 17, 2023, 3 pp. [cited by applicant]
Richiardi et al., “Speaker Verification With Confidence and Reliability Measures”, 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings., vol. 1, May 14, 2006, 1-4 pp. [cited by applicant]
Sanchez et al., “Multi-System Fusion of Extended Context Prosodic and Cepstral Features for Paralinguistic Speaker Trait Classification,” Thirteenth Annual Conference of the International Speech Communication Associatio… [cited by applicant]
Shriberg et al., “Effects of Vocal Effort and Speaking Style on Text-Independent Speaker Verification,” Ninth Annual Conference of the International Speech Communication Association, Sep. 2008, 4 pp. [cited by applicant]
Snyder, “NIST SRE 2016 Xvector Recipe”, Retrieved from: https://david-ryan-snyder.github.io/2017/10/04/model_sre16_v2.html, Oct. 4, 2017, 3 pp. [cited by applicant]
Talkin, “A Robust Algorithm for Pitch Tracking (RAPT)”, Speech Coding and Synthesis, Elsevier Science, 1995, pp. 495-502, (Applicant points out, in accordance with MPEP 609.04(a), that the year of publication, 1995, is … [cited by applicant]
Villalba et al., “Analysis of speech quality measures for the task of estimating the reliability of speaker verification decisions”, Speech Communication, vol. 78, Apr. 2016, 42-61 pp. [cited by applicant]
Villalba et al., “Reliability Estimation of the Speaker Verification Decisions Using Bayesian Networks to Combine Information from Multiple Speech Quality Measures”, Advances in Speech and Language Technologies for Iber… [cited by applicant]
Womack et al., “Classification of Speech Under Stress using Target Driven Features,” Speech Communications, Special Issue on Speech Under Stress, vol. 20, No. 1-2, Nov. 1996, pp. 131-150. [cited by applicant]
Womack et al., “N-Channel Hidden Markov Models for Combined Stress Speech Classification and Recognition,” IEEE Transactions on Speech & Audio Processing, vol. 7, No. 6, Nov. 1999, pp. 668-677. [cited by applicant]
Zhang et al., “An Advanced Entropy-based Feature with Frame-Level Vocal Effort Likelihood Space Modeling for Distant Whisper-Island Detection,” Speech Communication, vol. 66, Feb. 2015, pp. 107-117. [cited by applicant]
Zhang et al., “Analysis and Classification of Speech Mode: Whispered through Shouted,” Eighth Annual Conference of the International Speech Communication Association, Aug. 2007, 4 pp. [cited by applicant]
Zhang et al., “Whisper-Island Detection Based on Unsupervised Segmentation with Entropy-Based Speech Feature Processing,” IEEE Transactions Audio, Speech and Language Processing, vol. 19, No. 4, May 2011, pp. 883-894. [cited by applicant]
Zheng et al., “MobileUTDrive: An Android Portable Device Platform for In-vehicle Driving Data Collection and Display,” FAST-zero'15: 3rd International Symposium on Future Active Safety Technology Toward zero traffic acc… [cited by applicant]
Zhou et al., “Nonlinear Feature Based Classification of Speech under Stress,” IEEE Transactions on Speech & Audio Processing, vol. 9, No. 2, Mar. 2001, pp. 201-216. [cited by applicant]