IP Library Granted Patent US 11,636,872
Granted Patent B2
US 11,636,872 · App. 16/905,810 · Granted Apr 25, 2023

Techniques for computing perceived audio quality based on a trained multitask learning model

Inventors: Chih-Wei Wu (Los Gatos, CA); Phillip A. Williams (Los Gatos, CA); William Francis Wolcott, IV (Los Gatos, CA)
Assignee: NETFLIX, INC.
G10L25/60G06K9/623G06K9/6256G06N20/00G10L25/27G06F17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,872
App. No.
16/905,810
Granted
Apr 25, 2023
Kind
B2
Abstract

In various embodiments, a quality inference application estimates perceived audio quality. The quality inference application computes a set of feature values for a set of audio features based on an audio clip. The quality inference application then uses a trained multitask learning model to generate predicted labels based on the set of feature values. The predicted labels specify metric values for metrics that are relevant to audio quality. Subsequently, the quality inference application computes an audio quality score for the audio clip based on the predicted labels.

Claims (39)

1. A computer-implemented method for estimating perceived audio quality, the method comprising:

computing a first set of feature values for a set of audio features based on a first audio clip;

generating a first plurality of predicted labels via a trained multitask learning model based on the first set of feature values, wherein the first plurality of predicted labels specifies metric values for a plurality of metrics that are relevant to audio quality;

computing a geometric mean of the first plurality of predicted labels; and

scaling the geometric mean to generate a first audio quality score for the first audio clip.

2. The computer-implemented method of claim 1 , wherein scaling the geometric mean to generate computing the first audio quality score comprises scaling the geometric mean to a range of 1 to 5 corresponding to an Absolute Category Rating scale.

3. The computer-implemented method of claim 1 , wherein the first audio clip is derived from a reference audio clip using an audio algorithm, and wherein the first audio quality score indicates a perceptual impact associated with the audio algorithm.

4. The computer-implemented method of claim 1 , wherein generating the first plurality of predicted labels comprises inputting the first set of feature values into the trained multitask learning model that, in response, outputs the first plurality of predicted labels.

5. The computer-implemented method of claim 1 , wherein a first predicted label included in the first plurality of predicted labels comprises an estimated value of a scaled Hearing-Aid Audio Quality Index, a scaled Perceptual Evaluation of Audio Quality, a scaled Perception Model Quality, or a scaled Virtual Speech Quality Objective Listener Audio metric.

6. The computer-implemented method of claim 1 , wherein a first audio feature included in the set of audio features is associated with a first psycho-acoustic principle but not a second psycho-acoustic principle, and a second audio feature included in the set of audio features is associated with the second psycho-acoustic principle but not the first psycho-acoustic principle.

7. The computer-implemented method of claim 1 , wherein computing the first set of feature values for the set of audio features comprises:

computing a second set of feature values for a set of source features based on the first audio clip and a first reference clip; and

computing the first set of feature values for the set of audio features based on the second set of feature values for the set of source features and at least one scaling parameter associated with the trained multitask learning model.

8. The computer-implemented method of claim 1 , wherein the first audio clip includes at least one of dialogue, a sound effect, and background music.

9. The computer-implemented method of claim 1 , further comprising:

computing a second set of feature values for the set of audio features based on a second audio clip;

generating a second plurality of predicted labels via the trained multitask learning model based on the second set of feature values; and

computing a second audio quality score for the second audio clip based on the second plurality of predicted labels.

10. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to estimate perceived audio quality by performing the steps of:

computing a first set of feature values for a set of audio features based on a first audio clip;

causing a trained multitask learning model to generate a first plurality of predicted labels based on the first set of feature values, wherein the first plurality of predicted labels specifies metric values for a plurality of metrics that are relevant to audio quality;

computing a geometric mean of the first plurality of predicted labels; and

scaling the geometric mean to generate a first audio quality score for the first audio clip.

11. The one or more non-transitory computer readable media of claim 10 , wherein the first audio quality score is further generated by determining a plurality of weights associated with the first plurality of predicted labels based on one or more types of audio content included in the first audio clip.

12. The one or more non-transitory computer readable media of claim 10 , wherein the first audio clip includes at least one artifact associated with an audio algorithm, and wherein the first audio quality score indicates a quality versus processing efficiency tradeoff associated with the audio algorithm.

13. The one or more non-transitory computer readable media of claim 10 , wherein causing the trained multitask learning model to generate the first plurality of predicted labels comprises inputting the first set of feature values into the trained multitask learning model.

14. The one or more non-transitory computer readable media of claim 10 , wherein a first predicted label included in the first plurality of predicted labels comprises an estimated value of a scaled Hearing-Aid Audio Quality Index, a scaled Perceptual Evaluation of Audio Quality, a scaled Perception Model Quality, or a scaled Virtual Speech Quality Objective Listener Audio metric.

15. The one or more non-transitory computer readable media of claim 10 , wherein a first audio feature included in the set of audio features is associated with a first psycho-acoustic principle but not a second psycho-acoustic principle, and a second audio feature included in the set of audio features is associated with the second psycho-acoustic principle but not the first psycho-acoustic principle.

16. The one or more non-transitory computer readable media of claim 10 , wherein computing the first set of feature values comprises:

computing a first feature value for a first source feature based on the first audio clip and a first reference audio clip; and

performing at least one min-max scaling operation on the first feature value based on at least one scaling parameter associated with the trained multitask learning model to generate a second feature value for a first audio feature included in the set of audio features, wherein the second feature value is included in the first set of feature values.

17. The one or more non-transitory computer readable media of claim 10 , wherein the first audio clip includes at least one of dialogue, a sound effect, and background music.

18. A system comprising:

one or more memories storing instructions; and

one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:

computing a first set of feature values for a set of audio features based on a first audio clip;

generating a first plurality of predicted labels via a trained multitask learning model based on the first set of feature values, wherein the first plurality of predicted labels specifies metric values for a plurality of metrics that are relevant to audio quality;

performing at least one aggregation operation on the first plurality of predicted labels to compute an unscaled audio quality score for the first audio clip; and

scaling the unscaled audio quality score to generate an audio quality score for the first audio clip.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2020
From: WU, CHIH-WEI; WILLIAMS, PHILLIP A.; WOLCOTT, WILLIAM FRANCIS, IV
To: NETFLIX, INC.
Reel/Frame 053823/0147 →
Continuity (2)
Provisional Application 63021635 · May 7, 2020
Related Publication 20210350820A1 · Nov 11, 2021