IP Library › Granted Patent US 12,406,686
Granted Patent B2
US 12,406,686 · App. 16/905,793 · Granted Sep 2, 2025

Techniques for training a multitask learning model to assess perceived audio quality

Inventors: Chih-Wei Wu (Los Gatos, CA); Phillip A. Williams (Los Gatos, CA); William Francis Wolcott, IV (Los Gatos, CA)
Assignee: NETFLIX, INC.
G10L25/60G06F18/2113G06F18/214G06N20/00G10L25/27G06F17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,686
App. No.
16/905,793
Granted
Sep 2, 2025
Kind
B2
Abstract

In various embodiments, a training application trains a multitask learning model to assess perceived audio quality. The training application computes a set of pseudo labels based on a first audio clip and multiple models. The set of pseudo labels specifies metric values for a set of metrics that are relevant to audio quality. The training application also computes a set of feature values for a set of audio features based on the first audio clip. The training application trains a multitask learning model based on the set of feature values and the set of pseudo labels to generate a trained multitask learning model. In operation, the trained multitask learning model maps different sets of feature values for the set of audio features to different sets of predicted labels. Each set of predicted labels specifies estimated metric values for the set of metrics.

Claims (57)

1. A computer-implemented method for training a multitask learning model to assess perceived audio quality, the method comprising:

inputting a first training audio clip to a plurality of audio quality models that are separate from the multitask learning model to compute a first plurality of pseudo labels, wherein the first training audio clip is derived from a separate training reference audio clip, wherein each pseudo label included in the first plurality of pseudo labels specifies a metric value that is relevant to audio quality as measured by an audio quality model of the plurality of audio quality models;

performing one or more first scaling operations to scale each pseudo label included in the first plurality of pseudo labels based on a theoretical range of an output of a corresponding audio quality model included in the plurality of audio quality models to generate a first plurality of scaled pseudo labels;

computing a first set of feature values for a set of audio features based on the first training audio clip;

performing one or more second scaling operations to scale the first set of feature values using a set of scaling parameters computed based on the first set of feature values to generate a first set of scaled feature values;

training the multitask learning model based on the first set of scaled feature values and the first plurality of scaled pseudo labels to generate a trained multitask learning model that estimates the first plurality of scaled pseudo labels based on the first set of scaled feature values; and

transmitting the trained multitask learning model along with the set of scaling parameters to at least one compute instance, wherein the at least one compute instance, in operation, extracts a set of target features from a second audio clip, scales the set of target features based on the set of scaling parameters received with the trained multitask learning model to generate a second set of scaled feature values, and inputs the second set of scaled feature values to the trained multitask learning model, wherein the trained multitask learning model, in operation, maps, for the second audio clip, the second set of scaled feature values for the set of audio features based on the second audio clip to a plurality of predicted labels, and wherein the plurality of predicted labels specifies estimated metric values for a plurality of metrics relevant to audio quality that would be computed according to the plurality of audio quality models for the second audio clip.

2. The computer-implemented method of claim 1 , wherein a first audio quality model included in the plurality of audio quality models comprises a Hearing-Aid Audio Quality Index expert system, a Perceptual Evaluation of Audio Quality expert system, a Perception Model Quality expert system, or a Virtual Speech Quality Objective Listener Audio expert system.

3. The computer-implemented method of claim 1 , wherein a first audio feature included in the set of audio features based on the first training audio clip is associated with a Cepstral Correlation, a Noise-to-Mask Ratio, a Perceptual Similarity Measure, a Neurogram Similarity Index Measure, or a bitrate.

4. The computer-implemented method of claim 1 , wherein computing the first set of feature values comprises:

computing a plurality of source feature values for a first source feature based on a plurality of training audio clips that includes the first training audio clip;

computing at least one scaling parameter included in the set of scaling parameters based on the plurality of source feature values; and

computing a first feature value for a first audio feature included in the set of audio features based on a first source feature value included in the plurality of source feature values and the at least one scaling parameter, wherein the first feature value is included in the first set of feature values.

5. The computer-implemented method of claim 4 , wherein computing the at least one scaling parameter comprises:

setting a first scaling parameter equal to a minimum source feature value included in the plurality of source feature values; and

setting a second scaling parameter equal to a maximum source feature value included in the plurality of source feature values.

6. The computer-implemented method of claim 4 , further comprising computing an overall quality score for the second audio clip based on the trained multitask learning model and the at least one scaling parameter.

7. The computer-implemented method of claim 1 , wherein:

computing a first pseudo label included in the first plurality of pseudo labels comprises inputting the first training audio clip and the separate training reference audio clip into a first audio quality model included in the plurality of audio quality models that, in response, outputs a first metric value for a first metric included in the plurality of metrics, and

scaling the first pseudo label to generate a first scaled pseudo label included in the first plurality of scaled pseudo labels comprises scaling the first metric value based on a theoretical minimum and a theoretical maximum of the first metric computed by the first audio quality model.

8. The computer-implemented method of claim 1 , wherein training the multitask learning model based on the first set of scaled feature values and the first plurality of scaled pseudo labels comprises:

inputting the first set of scaled feature values into the multitask learning model that, in response, outputs a first plurality of predicted labels;

computing a loss based on the first plurality of predicted labels and the first plurality of scaled pseudo labels; and

performing one or more optimization operations on the multitask learning model based on the loss.

9. The computer-implemented method of claim 8 , wherein computing the loss comprises computing a mean squared error between the first plurality of predicted labels and the first plurality of scaled pseudo labels.

10. The computer-implemented method of claim 1 , wherein training the multitask learning model based on the first set of scaled feature values and the first plurality of pseudo labels comprises executing a multitask learning algorithm on the multitask learning model based on the first set of scaled feature values, the first plurality of scaled pseudo labels, an additional set of feature values associated with an additional training audio clip, and an additional plurality of scaled feature values associated with the additional training audio clip.

11. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to train a multitask learning model to assess perceived audio quality by performing the steps of:

inputting a first training audio clip to a plurality of audio quality models that are separate from the multitask learning model to generate a plurality of output values, wherein the first training audio clip is derived from a separate training reference audio clip;

computing a plurality of pseudo labels based on the plurality of output values, wherein each pseudo label included in the plurality of pseudo labels specifies a metric value that is relevant to audio quality as measured by an audio quality model of the plurality of audio quality models;

performing one or more first scaling operations to scale each pseudo label included in the plurality of pseudo labels based on a theoretical range of an output of a corresponding audio quality model included in the plurality of audio quality models to generate a plurality of scaled pseudo labels;

computing a first set of feature values for a set of audio features based on the first training audio clip;

performing one or more second scaling operations to scale the first set of feature values using a set of scaling parameters computed based on the first set of feature values to generate a first set of scaled feature values;

training the multitask learning model based on the first set of scaled feature values and the plurality of scaled pseudo labels to generate a trained multitask learning model that estimates the plurality of scaled pseudo labels based on the first set of scaled feature values; and

transmitting the trained multitask learning model along with the set of scaling parameters to at least one compute instance, wherein the at least one compute instance, in operation, extracts a set of target features from a second audio clip, scales the set of target features based on the set of scaling parameters received with the trained multitask learning model to generate a second set of scaled feature values, and inputs the second set of scaled feature values to the trained multitask learning model, wherein the trained multitask learning model, in operation, maps, for the second audio clip, the second set of scaled feature values for the set of audio features based on the second audio clip to a plurality of predicted labels, and wherein the plurality of predicted labels specifies estimated metric values for a plurality of metrics relevant to audio quality that would be computed according to the plurality of audio quality models for the second audio clip.

12. The one or more non-transitory computer readable media of claim 11 , wherein a first audio quality model included in the plurality of audio quality models comprises a perceptual quality model that is trained based on subjective scores assigned by human listeners.

13. The one or more non-transitory computer readable media of claim 11 , wherein a first audio feature included in the set of audio features is associated with a Cepstral Correlation, a Noise-to-Mask Ratio, a Perceptual Similarity Measure, a Neurogram Similarity Index Measure, or a bitrate.

14. The one or more non-transitory computer readable media of claim 11 , wherein the multitask learning model comprises at least one of a neural network, a decision tree, or a random forest.

15. The one or more non-transitory computer readable media of claim 11 , wherein computing the first set of feature values comprises:

computing a plurality of source feature values for a first source feature based on a plurality of training audio clips that includes the first training audio clip;

computing at least one scaling parameter included in the set of scaling parameters based on the plurality of source feature values; and

computing a first feature value for a first audio feature included in the set of audio features based on a first source feature value included in the plurality of source feature values and the at least one scaling parameter, wherein the first feature value is included in the first set of feature values.

16. The one or more non-transitory computer readable media of claim 15 , wherein computing the first feature value comprises performing at least one min-max scaling operation on the first source feature value.

17. The one or more non-transitory computer readable media of claim 15 , further comprising computing an overall quality score for the second audio clip based on the trained multitask learning model and the at least one scaling parameter.

18. The one or more non-transitory computer readable media of claim 11 , wherein a first pseudo label included in the plurality of pseudo labels is scaled based on a theoretical minimum and a theoretical maximum of a first output of a first audio quality model included in the plurality of audio quality models.

19. The one or more non-transitory computer readable media of claim 11 , wherein training the multitask learning model based on the first set of scaled feature values and the plurality of scaled pseudo labels comprises:

inputting the first set of scaled feature values into the multitask learning model that, in response, outputs a first plurality of predicted labels;

computing a loss based on the first plurality of predicted labels and the plurality of scaled pseudo labels; and

performing one or more optimization operations on the multitask learning model based on the loss.

20. A system comprising:

one or more memories storing instructions; and

one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:

inputting a first training audio clip to a plurality of audio quality models that are separate from the multitask learning model to compute a first plurality of pseudo labels, wherein the first training audio clip is derived from a separate training reference audio clip, wherein each pseudo label included in the first plurality of pseudo labels specifies a metric value that is relevant to audio quality as measured by an audio quality model of the plurality of audio quality models;

performing one or more first scaling operations to scale each pseudo label included in the first plurality of pseudo labels based on a theoretical range of an output of a corresponding audio quality model included in the plurality of audio quality models to generate a first plurality of scaled pseudo labels;

computing a first set of feature values for a set of audio features based on the first training audio clip;

performing one or more second scaling operations to scale the first set of feature values using a set of scaling parameters computed based on the first set of feature values to generate a first set of scaled feature values;

executing at least one multitask learning algorithm based on the first set of scaled feature values and the first plurality of scaled pseudo labels to generate a trained multitask learning model that estimates the first plurality of scaled pseudo labels based on the first set of scaled feature values; and

transmitting the trained multitask learning model along with the set of scaling parameters to at least one compute instance, wherein the at least one compute instance, in operation, extracts a set of target features from a second audio clip, scales the set of target features based on the set of scaling parameters received with the trained multitask learning model to generate a second set of scaled feature values, and inputs the second set of scaled feature values to the trained multitask learning model, wherein the trained multitask learning model, in operation, maps, for the second audio clip, the second set of scaled feature values for the set of audio features based on the second audio clip to a plurality of predicted labels, and wherein the plurality of predicted labels specifies estimated metric values for a plurality of metrics relevant to audio quality that would be computed according to the plurality of audio quality models for the second audio clip.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2020
From: WU, CHIH-WEI; WILLIAMS, PHILLIP A.; WOLCOTT, WILLIAM FRANCIS, IV
To: NETFLIX, INC.
Reel/Frame 053823/0144 →
Continuity (2)
Provisional Application 63021635 · May 7, 2020
Related Publication 20210350819A1 · Nov 11, 2021
References Cited (83)
US 8467893B2 · Grancharov · 2013 [cited by examiner]
US 9025780B2 · Beerends · 2015 [cited by examiner]
US 9396738B2 · Abdelal · 2016 [cited by examiner]
US 9635483B2 · Francombe · 2017 [cited by examiner]
US 9870784B2 · Sharma · 2018 [cited by examiner]
US 10049674B2 · Xiao · 2018 [cited by examiner]
US 10667155B2 · Ouyang · 2020 [cited by examiner]
US 10957337B2 · Chen · 2021 [cited by examiner]
US 10984818B2 · Xiao · 2021 [cited by examiner]
US 20180158470A1 · Zhu et al. · 2018 [cited by applicant]
US 20190385480A1 · Suzuki · 2019 [cited by examiner]
US 20200022007A1 · Ouyang et al. · 2020 [cited by applicant]
US 20200227070A1 · Kim et al. · 2020 [cited by applicant]
US 20200349467A1 · Teague · 2020 [cited by examiner]
US 20200402530A1 · Güzelarslan · 2020 [cited by examiner]
US 20210125629A1 · Bryan · 2021 [cited by examiner]
US 20210217403A1 · Chae · 2021 [cited by applicant]
US 20210256988A1 · Mauri et al. · 2021 [cited by applicant]
US 20210264938A1 · Yao et al. · 2021 [cited by applicant]
US 20210272583A1 · Liu · 2021 [cited by applicant]
US 20210350820A1 · Wu · 2021 [cited by examiner]
Kang, Myeongsu, and Jing Tian, “Machine Learning: Data Pre-processing”, 2018, Prognostics and Health Management of Electronics: Fundamentals, Machine Learning, and the Internet of Things, Chapter 5, pp. 111-130. (Year: … [cited by examiner]
Dong, Xuan, and Donald S. Williamson, “An Attention Enhanced Multi-Task Model for Objective Speech Assessment in Real-World Environments”, Apr. 2020, ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech a… [cited by examiner]
Sharma, Dushyant, Aidan O. T. Hogg, Yu Wang, Amr H. Nour-Eldin, and Patrick A. Naylor, “Non-Intrusive POLQA Estimation of Speech Quality using Recurrent Neural Networks”, 2019, 2019 27th European Signal Processing Confe… [cited by examiner]
International Search Report for application No. PCT/US2021/030968 dated Jul. 30, 2021. [cited by applicant]
Choi et al., “Neural MOS prediction for synthesized speech using multi-task learning with spoofing detection and spoofing type classification”, 2021 IEEE spoken language technology workshop (SLT), IEEE, DOI:10.1109/SLT4… [cited by applicant]
Sharma et al., “Non-Intrusive POLQA Estimation of Speech Quality using Recurrent Neural Networks”, 27th European Signal Processing Conference (EUSIPCO), DOI: 10.23919/EUSIPC0.2019.8902646, Sep. 2, 2019, 5 bages. [cited by applicant]
Dong et al., “An Attention Enhanced Multi-Task Model for Objective Speech Assessment in Real-World Environments”, IEEE International Conference on Acoustics, Speech and Signal processing (ICASSP), IEEE, DOI: 10.1109/ICA… [cited by applicant]
Avila et al., “Non-Intrusive Speech Quality Assessment Using Neural Networks”, In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Mar. 2019, pp. 1-5. [cited by applicant]
Beerends et al., “Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I—Temporal Alignment”, Journal of the Audio Engineering Soc… [cited by applicant]
Biberger et al., “An Objective Audio Quality Measure Based on Power and Envelope Power Cues”, Journal of the Audio Engineering Society, DOI: https://doi.org/10.17743/jaes.2018.0031, vol. 66, No. 7/8, Jul./Aug. 2018, pp.… [cited by applicant]
Bock et al., “Multi-Task Learning of Tempo and Beat: Learning One To Improve the Other”, In Proceedings of the 20th ISMIR Conference, Nov. 4-8, 2019, pp. 486-493. [cited by applicant]
Campbell et al., “Audio quality assessment techniques—A review, and recent developments”, Signal Processing, doi:10.1016/j.sigpro.2009.02.015, vol. 89, No. 8, Aug. 2009, pp. 1489-1500. [cited by applicant]
Caruana, Rich, “Multitask Learning”, Machine learning, vol. 28, No. 35, 1997, pp. 41-75. [cited by applicant]
Chen et al., “Functional Harmony Recognition of Symbolic Music Data With Multi-Task Recurrent Neural Networks”, In Proceedings of the 19th ISMIR Conference, Sep. 23-27, 2018, pp. 90-97. [cited by applicant]
Dau et al., “Modeling auditory processing of amplitude modulation. I. Detection and masking with narrow-band carriers”, The Journal of the Acoustical Society of America, vol. 102, No. 5, Nov. 1997, pp. 2892-2905. [cited by applicant]
Dau et al., “Modeling auditory processing of amplitude modulation. II. Spectral and temporal integration”, The Journal of the Acoustical Society of America, vol. 102, No. 5, Nov. 1997, pp. 2906-2919. [cited by applicant]
Emiya et al., “Subjective and Objective Quality Assessment of Audio Source Separation”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, No. 7, Sep. 2011, pp. 2046-2057. [cited by applicant]
Hines et al., “VISQOLAudio: An objective audio quality metric for low bitrate codecs”, Journal of the Acoustical Society of America, http://dx.doi.org/10.1121/1.4921674, vol. 137, No. 6, Jun. 2015, pp. EL449-EL455. [cited by applicant]
Hines et al., “VISQOL: an objective speech quality model”, EURASIP Journal on Audio, Speech, and Music Processing, DOI 10.1186/s13636-015-0054-9, vol. 1, No. 13, Dec. 2015, pp. 1-18. [cited by applicant]
Huber et al., “PEMO-Q—A New Method for Objective Audio Quality Assessment Using a Model of Auditory Perception”, IEEE Transactions on Audio, Speech and Language Processing, vol. 14, No. 6, Nov. 2006, pp. 1902-1911. [cited by applicant]
Hung et al., “Multitask Learning for Frame-Level Instrument Recognition”, In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing (ICASSP), arXiv:1811.01143, 2019, 5 pages. [cited by applicant]
Kates et al., “The Hearing-Aid Speech Quality Index (HASQI) Version 2”, Journal of the Audio Engineering Society, vol. 62, No. 3, Mar. 2014, pp. 99-117. [cited by applicant]
Kates et al., “The Hearing-Aid Audio Quality Index (Haaqi)”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, No. 2, Feb. 2016, pp. 354-365. [cited by applicant]
Manocha et al., “A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences”, In arXiv—arXiv:2001.04460, May 2020, 6 pages. [cited by applicant]
Martinez et al., “A No-Reference Audio-Visual Video Quality Metric”, In Proceedings of the European Signal Processing Conference (EUSIPCO), 2014, 5 pages. [cited by applicant]
Moore, Brian C. J., “Computational models for predicting sound quality”, Acoustical Science and Technology, vol. 41, No. 1, Jan. 2020, pp. 75-82. [cited by applicant]
Pocta et al., “Subjective and Objective Assessment of Perceived Audio Quality of Current Digital Audio Broadcasting Systems and Web-Casting Applications”, IEEE Transactions on Broadcasting, vol. 61, No. 3, Sep. 2015, pp… [cited by applicant]
Rix et al., “Perceptual Evaluation of Speech Quality (PESQ)—A New Method for Speech Quality Assessment of Telephone Networks and Codecs”, In Proceedings of the International Conference on Acoustics, Speech, and Signal P… [cited by applicant]
Rix et al., “Objective Assessment of Speech and Audio Quality—Technology and Applications”, IEEE Transactions on Audio, Speech and Language Processing, DOI 10.1109/TASL.2006.883260, vol. 14, No. 6, Nov. 2006, pp. 1890-1… [cited by applicant]
Sloan et al., “Objective Assessment of Perceptual Audio Quality Using VISQOLAudio”, IEEE Transactions on Broadcasting, DOI 10.1109/TBC.2017.2704421, vol. 63, No. 4, Dec. 2017, pp. 693-705. [cited by applicant]
Thiede et al., “PEAQ—The ITU Standard for Objective Measurement of Perceived Audio Quality”, Journal of the Audio Engineering Society, vol. 48, No. 1, 2000, pp. 3-29. [cited by applicant]
Torcoli et al., “Comparing the Effect of Audio Coding Artifacts on Objective Quality Measures and on Subjective Ratings”, In Proceedings of the Audio Engineering Society (AES) convention, May 23-26, 2018, pp. 1-10. [cited by applicant]
International Telecommunication Union, Recommendation ITU-T P.861, “Objective quality measurement of telephone-band (300-3400 Hz) speech codecs”, 1996, 34 pages. [cited by applicant]
International Telecommunication Union, Recommendation ITU-T P.563:, “Single-ended method for objective speech quality assessment in narrow-band telephony applications”, 2004, 66 pages. [cited by applicant]
International Telecommunication Union, Recommendation ITU-R BS.1116-3:, “Methods for the subjective assessment of small impairments in audio systems”, 2015, pp. 1-32. [cited by applicant]
International Telecommunication Union, Recommendation ITU-R BS. 1534-3:, “Method for the subjective assessment of intermediate quality level of audio systems”, 2015, pp. 1-36. [cited by applicant]
Wang et al., “Image Quality Assessment: From Error Visibility to Structural Similarity”, IEEE Transactions on Image Processing, vol. 13, No. 4, Apr. 2004, pp. 600-612. [cited by applicant]
Zielinski et al., “On Some Biases Encountered in Modern Audio Quality Listening Tests—A Review”, Journal of the Audio Engineering Society, vol. 56, No. 6, Jun. 2008, pp. 427-451. [cited by applicant]
TSP Lab Software, “Telecommunications & Signal Processing Laboratory Multimedia Signal Processing”, Retrieved from http://www-mmsp.ece.mcgill.ca/Documents/Software, on Jan. 21, 2021, 3 pages. [cited by applicant]
The PEASS Toolkit—Perceptual Evaluation methods for Audio Source Separation, “The PEASS Software”, Retrieved from http://bass-db.gforge.inria.fr/peass/PEASS-Software.html, on Jan. 21, 2021, 2 pages. [cited by applicant]
Audio and Music, “Quality of Experience Research Lab | QxLab”, Retrieved from https://qxlab.ucd.ie/index.php/audio-and-music, on Jan. 14, 2021, 2 pages. [cited by applicant]
Librosa, Retrieved from http://librosa.org/doc/latest/index.html, on Jan. 21, 2021, 3 pages. [cited by applicant]
GitHub, “alexanderlerch/pyACA”, Retrieved from https://github.com/alexanderlerch/pyACA, on Jan. 21, 2021, 4 pages. [cited by applicant]
GitHub, scikit-leam, “Machine Learning in Python”, Release Highlights for 0.24, Retrieved from https://scikit-learn.org/stable/, 2021, 2 pages. [cited by applicant]
TensorFlow Core v2.4.0, “Module: tf.keras”, Retrieved from https://www.tensorflow.org/api_docs/python/tf/keras, on Jan. 21, 2021, 3 pages. [cited by applicant]
FFmpeg, “A complete, cross-platform solution to record, convert and stream audio and video”, Retrieved from https://www.ffmpeg.org/, on Jan. 21, 2021, 28 pages. [cited by applicant]
International Telecommunication Union, Recommendation ITU-R BS. 1387-1, Retrieved from https://www.itu.int/rec/R-REC-BS.1387-1-200111-l/en, 1998-2001, pp. 1-100. [cited by applicant]
Results of the public multiformat listening test, Retrieved from http://listening-test.coresv.net/results.htm, on Jan. 21, 2021, 49 pages. [cited by applicant]
Video Quality Databases, Retrieved from http://www.ene.unb.br/mylene/databases.html, on Jan. 21, 2021, 6 pages. [cited by applicant]
GitHub, “Perceptual-Coding-In-Python”, Retrieved from https://github.com/stephencwelch/Perceptual-Coding-In-Python/tree/master/PEAQPython/PQevalAudioMATLAB, on Jan. 21, 2021, 2 pages. [cited by applicant]
GitHub, “Perceptual-Coding-In-Python”, Retrieved from https://github.com/stephencwelch/Perceptual-Coding-In-Python/tree/master/PEAQPython, on Jan. 21, 2021, 1 page. [cited by applicant]
GitHub, “EAQUAL”, Retrieved from https://github.com/godock/eaqual, on Jan. 21, 2021, 5 pages. [cited by applicant]
GitHub, “python-pesq”, Retrieved from https://github.com/vBaiCai/python-pesq, on Jan. 21, 2021, 4 pages. [cited by applicant]
GitHub, “mos-ivr”, Retrieved from https://github.com/stingerpk/mos-ivr/tree/master/P563, on Jan. 21, 2021, 5 pages. [cited by applicant]
GitHub, “peass-software”, Retrieved from https://github.com/CVSSP/peass-software/blob/master/pemo_metric.m, on Jan. 21, 2021, 3 pages. [cited by applicant]
GitHub, “dmca”, Retrieved from https://github.com/github/dmca/blob/master/2018/2018-02-14-POLQA.md, on Jan. 21, 2021, 5 pages. [cited by applicant]
Non Final Office Action received for U.S. Appl. No. 16/905,810 dated Apr. 1, 2022, 32 pages. [cited by applicant]
Wilkinghoff et al., “Robust Speaker Identification by Fusing Classification Scores with a Neural Network”, Speech Communication; 13th ITG-Symposium, VDE, Oct. 2018, pp. 261-265. [cited by applicant]
Zhang et al., “Unsupervised Learning in Cross-Corpus Acoustic Emotion Recognition”, 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, IEEE, 2011, pp. 523-528. [cited by applicant]
Olsan et al., “Geometric Mean Technique”, Decision Aids for Selection Problems, 1996, pp. 69-80. [cited by applicant]
Final Office Action received for U.S. Appl. No. 16/905,810 dated Jul. 20, 2022, 20 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 16/905,810 dated Dec. 7, 2022, 14 pages. [cited by applicant]