Techniques for limiting the influence of image enhancement operations on perceptual video quality estimations
In various embodiments, a tunable VMAF application reduces an amount of influence that image enhancement operations have on perceptual video quality estimates. In operation, the tunable VMAF application computes a first value for a first visual quality metric based on reconstructed video content and a first enhancement gain limit. The tunable VMAF application computes a second value for a second visual quality metric based on the reconstructed video content and a second enhancement gain limit. Subsequently, the tunable VMAF application generates a feature value vector based on the first value for the first visual quality metric and the second value for the second visual quality metric. The tunable VMAF application executes a VMAF model based on the feature value vector to generate a tuned VMAF score that accounts, at least in part, for at least one image enhancement operation used to generate the reconstructed video content.
1. A computer-implemented method for reducing the impact that image enhancement operations have on perceptual video quality estimates, the method comprising:
computing a first value for a first visual quality metric based on first reconstructed video content, wherein a first image enhancement operation is used to generate the first reconstructed video content;
computing a second value for a second visual quality metric based on the first reconstructed video content;
generating a first feature value vector based on the first value and the second value; and
generating, based on the first feature value vector, a first tuned Video Multi-Method Assessment Fusion (VMAF) score that accounts, at least in part, for the first image enhancement operation.
2. The computer-implemented method of claim 1 , wherein computing the first value for the first visual quality metric is further based on a first enhancement gain limit.
3. The computer-implemented method of claim 1 , wherein computing the second vale for the second visual quality metric is further based on a second enhancement gain limit.
4. The computer-implemented method of claim 1 , wherein generating the first turned VMAF score comprises executing a VMAF model based on the first feature value vector.
5. The computer-implemented method of claim 4 , wherein the VMAF model implements at least one of a support vector regression algorithm, an artificial neural network algorithm, or a random forest algorithm that is trained based on a plurality of human-observed visual quality scores for reconstructed training video content.
6. The computer-implemented method of claim 1 , wherein the first tuned VMAF score reduces how much impact the at least one image enhancement operation has on a perceptual video quality associated with the first reconstructed video content.
7. The computer-implemented method of claim 1 , wherein the second visual quality metric comprises a modified version of a wavelet-domain implementation of a detail loss metric (“DLM”).
8. The computer-implemented method of claim 1 , wherein the first value for the first visual quality metric is associated with a first spatial scale.
9. The computer-implemented method of claim 1 , wherein the first reconstructed video content comprises a first frame included in a reconstructed video, and further comprising computing an overall tuned VMAF score associated with the reconstructed video based on the first tuned VMAF score and at least a second tuned VMAF score associated with a second frame included in the reconstructed video.
10. The computer-implemented method of claim 1 , wherein the first image enhancement operation comprises a sharpening operation, a contrasting operation, or a histogram equalization operation.
11. The computer-implemented method of claim 1 , further comprising generating the first reconstructed video content using a coder/decoder (“codec”) that executes a first image enhancement operation either prior to executing at least one data compression operation or subsequent to executing at least one data decompression operation.
12. One or more non-transitory computer-readable media including instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:
generating a first feature value vector based on at least one value of a visual quality metric associated with reconstructed video content; and
generating, based on the first feature value vector, a first tuned Video Multi-Method Assessment Fusion (VMAF) score that accounts, at least in part, for a first image enhancement operation.
13. The one or more non-transitory computer-readable media of claim 12 , further comprising computing the first value for the first visual quality metric based on the first reconstructed video content.
14. The one or more non-transitory computer-readable media of claim 13 , wherein computing the first value for the first visual quality metric is further based on a first enhancement gain limit.
15. The one or more non-transitory computer-readable media of claim 12 , wherein the first image enhancement operation is used to generate the first reconstructed video content.
16. The one or more non-transitory computer-readable media of claim 15 , wherein the first image enhancement operation comprises a sharpening operation, a contrasting operation, or a histogram equalization operation.
17. The one or more non-transitory computer-readable media of claim 12 , further comprising computing a second value for a second visual quality metric based on the first reconstructed video content.
18. The one or more non-transitory computer-readable media of claim 17 , wherein computing the second vale for the second visual quality metric is further based on a second enhancement gain limit.
19. The one or more non-transitory computer readable media of claim 12 , wherein the first visual quality metric comprises a modified version of a pixel-domain implementation of a Visual Information Fidelity (“VIF”) quality index.
20. The one or more non-transitory computer readable media of claim 12 , wherein the first reconstructed video content comprises a first frame included in a reconstructed video, and further comprising performing one or more temporal pooling operations on a plurality of tuned VMAF scores that includes the first tuned VMAF score and is associated with a plurality of frames included in the reconstructed video to determine an overall tuned VMAF score associated with the reconstructed video.
21. The one or more non-transitory computer-readable media of claim 12 , wherein generating the first turned VMAF score comprises executing a VMAF model based on the first feature value vector.
22. The one or more non-transitory computer-readable media of claim 12 , wherein the VMAF model implements at least one of a support vector regression algorithm, an artificial neural network algorithm, or a random forest algorithm that is trained based on a plurality of human-observed visual quality scores for reconstructed training video content.
23. The one or more non-transitory computer-readable media of claim 12 , further comprising generating the first reconstructed video content using a coder/decoder (“codec”) that executes a first image enhancement operation either prior to executing at least one data compression operation or subsequent to executing at least one data decompression operation.
24. A system comprising:
one or more memories storing instructions; and
one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:
generating a first feature value vector based on at least one value of a visual quality metric; and
generating, based on the first feature value vector, a first tuned Video Multi-Method Assessment Fusion (VMAF) score based on a first image enhancement operation used to generate reconstructed video content.