IP Library Granted Patent US 12,418,685
Granted Patent B2
US 12,418,685 · App. 17/549,793 · Granted Sep 16, 2025

Techniques for limiting the influence of image enhancement operations on perceptual video quality estimations

Inventor: Zhi Li (Mountain View, CA)
Assignee: NETFLIX, INC.
H04N19/86
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,418,685
App. No.
17/549,793
Granted
Sep 16, 2025
Kind
B2
Abstract

In various embodiments, a tunable VMAF application reduces an amount of influence that image enhancement operations have on perceptual video quality estimates. In operation, the tunable VMAF application computes a first value for a first visual quality metric based on reconstructed video content and a first enhancement gain limit. The tunable VMAF application computes a second value for a second visual quality metric based on the reconstructed video content and a second enhancement gain limit. Subsequently, the tunable VMAF application generates a feature value vector based on the first value for the first visual quality metric and the second value for the second visual quality metric. The tunable VMAF application executes a VMAF model based on the feature value vector to generate a tuned VMAF score that accounts, at least in part, for at least one image enhancement operation used to generate the reconstructed video content.

Claims (41)

1. A computer-implemented method, comprising:

computing a first value for a first visual quality metric for first reconstructed video content;

computing a second value for a second visual quality metric for the first reconstructed video content;

generating a first feature value vector based on the first value and the second value; and

generating, based on the first feature value vector, a score that accounts for, at least in part, a first image enhancement operation used to generate the first reconstructed video content.

2. The computer-implemented method of claim 1 , wherein computing the first value for the first visual quality metric is based on a first enhancement gain limit.

3. The computer-implemented method of claim 1 , wherein computing the second value for the second visual quality metric is based on a second enhancement gain limit.

4. The computer-implemented method of claim 1 , wherein generating the score comprises executing a trained model based on the first feature value vector.

5. The computer-implemented method of claim 4 , wherein the trained model implements at least one of a support vector regression algorithm, an artificial neural network algorithm, or a random forest algorithm that is trained based on a plurality of human-observed visual quality scores for reconstructed training video content.

6. The computer-implemented method of claim 1 , wherein the score reduces how much impact the first image enhancement operation has on a perceptual video quality associated with the first reconstructed video content.

7. The computer-implemented method of claim 1 , wherein the second visual quality metric comprises a modified version of a wavelet-domain implementation of a detail loss metric (“DLM”).

8. The computer-implemented method of claim 1 , wherein the first value for the first visual quality metric is associated with a first spatial scale.

9. The computer-implemented method of claim 1 , wherein the first value for the first visual quality metric is computed based on the first reconstructed video content that comprises a first frame included in a reconstructed video, and further comprising computing an overall score for the reconstructed video based on the score and at least a second score for a second frame included in the reconstructed video.

10. The computer-implemented method of claim 1 , wherein the first image enhancement operation comprises a sharpening operation, a contrasting operation, or a histogram equalization operation.

11. The computer-implemented method of claim 1 , further comprising generating the first reconstructed video content using a coder/decoder (“codec”) that executes the first image enhancement operation.

12. One or more non-transitory computer-readable media including instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

generating a first feature value vector based on a first value for a first visual quality metric for first reconstructed video content and a second value for a second visual quality metric for the first reconstructed video content; and

generating, based on the first feature value vector, a first score that accounts for, at least in part, a first image enhancement operation used to generate the first reconstructed video content.

13. The one or more non-transitory computer-readable media of claim 12 , further comprising computing the first value for the first visual quality metric based on a first enhancement gain limit.

14. The one or more non-transitory computer-readable media of claim 12 , wherein the first image enhancement operation comprises a sharpening operation, a contrasting operation, or a histogram equalization operation.

15. The one or more non-transitory computer-readable media of claim 12 , further comprising computing the second value for the second visual quality metric based on a second enhancement gain limit.

16. The one or more non-transitory computer-readable media of claim 12 , wherein the first visual quality metric comprises a modified version of a pixel-domain implementation of a Visual Information Fidelity (“VIF”) quality index.

17. The one or more non-transitory computer-readable media of claim 12 , wherein the first reconstructed video content comprises a first frame included in a reconstructed video, and further comprising performing one or more temporal pooling operations on a plurality of scores that includes the first score and are for a plurality of frames included in the reconstructed video to determine an overall score for the reconstructed video.

18. The one or more non-transitory computer-readable media of claim 12 , wherein generating the first score comprises executing a trained model based on the first feature value vector.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the trained model implements at least one of a support vector regression algorithm, an artificial neural network algorithm, or a random forest algorithm that is trained based on a plurality of human-observed visual quality scores for reconstructed training video content.

20. The one or more non-transitory computer-readable media of claim 12 , further comprising generating the first reconstructed video content using a coder/decoder (“codec”) that executes the first image enhancement operation.

21. A system comprising:

one or more memories storing instructions; and

one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:

generating a first feature value vector based on at least a first value of a first visual quality metric for first reconstructed video content and a second value for a second visual quality metric for the first reconstructed video content; and

generating, based on the first feature value vector, a score that accounts for, at least in part, a first image enhancement operation used to generate the first reconstructed video content.

22. A computer-implemented method for reducing the impact that image enhancement operations have on perceptual video quality estimates, the method comprising:

computing a first value for a first visual quality metric based on first reconstructed video content and a first enhancement gain limit, wherein a first image enhancement operation is used to generate the first reconstructed video content;

computing a second value for a second visual quality metric based on the first reconstructed video content and a second enhancement gain limit;

generating a first feature value vector based on the first value and the second value; and

generating, based on the first feature value vector, a first score that accounts for, at least in part, the first image enhancement operation.

23. The computer-implemented method of claim 22 , wherein the first score comprises a first tuned Video Multi-Method Assessment Fusion (VMAF) score.

24. The computer-implemented method of claim 22 , wherein generating the first score comprises executing a trained model based on the first feature value vector.

25. The computer-implemented method of claim 24 , wherein the trained model implements at least one of a support vector regression algorithm, an artificial neural network algorithm, or a random forest algorithm that is trained based on a plurality of human-observed visual quality scores for reconstructed training video content.

26. The computer-implemented method of claim 22 , wherein the first reconstructed video content comprises a first frame included in a reconstructed video, and further comprising computing an overall score for the reconstructed video based on the first score and at least a second score for a second frame included in the reconstructed video.

27. The computer-implemented method of claim 22 , wherein the first image enhancement operation comprises a sharpening operation, a contrasting operation, or a histogram equalization operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: LI, ZHI
To: NETFLIX, INC.
Reel/Frame 059075/0941 →
Continuity (4)
Continuation 17157871 · Jan 25, 2021
Provisional Application 63117931 · Nov 24, 2020
Provisional Application 63052423 · Jul 15, 2020
Related Publication 20220103869A1 · Mar 31, 2022
References Cited (24)
US 7526142B2 · Sheraizin et al. · 2009 [cited by applicant]
US 10255667B2 · He et al. · 2019 [cited by applicant]
US 10475172B2 · Aaron et al. · 2019 [cited by applicant]
US 10715814B2 · Katsavounidis · 2020 [cited by applicant]
US 11089359B1 · Katsavounidis · 2021 [cited by applicant]
US 20110282194A1 · Reiner · 2011 [cited by examiner]
US 20160021380A1 · Li et al. · 2016 [cited by applicant]
US 20160227220A1 · Tanchenko · 2016 [cited by examiner]
US 20160317127A1 · dos Santos Mendonca · 2016 [cited by examiner]
US 20160335754A1 · Aaron · 2016 [cited by examiner]
US 20170132785A1 · Wshah · 2017 [cited by examiner]
US 20180144214A1 · Hsieh · 2018 [cited by examiner]
US 20190246112A1 · Li et al. · 2019 [cited by applicant]
US 20190289296A1 · Kottke et al. · 2019 [cited by applicant]
US 20190373293A1 · Bortman et al. · 2019 [cited by applicant]
US 20200021865A1 · Topiwala · 2020 [cited by examiner]
US 20200145661A1 · Jeon et al. · 2020 [cited by applicant]
US 20200296362A1 · Chadwick et al. · 2020 [cited by applicant]
US 20210385502A1 · Dinh · 2021 [cited by examiner]
WO 2020043279A1 · 2020 [cited by applicant]
Moldonvan et al., “A Novel Mechanism for Mapping Objective Video Quality Metrics to Subjective MOS Scale”, DOI:10.1109/BMSB.2014.6873572, Conference: 2014 IEEE International Symposium on Broadband Multimedia Systems and… [cited by applicant]
Bampis et al., “Enhancing Temporal Quality Measurements in a Globally Deployed Streaming Video Quality Predictor,” 2018 25th IEEE International Conference on Imaging Processing (ICIP), Athens, Greece, 2018, pp. 614-618,… [cited by applicant]
Li et al., “Image Quality Assesment by Separately Evaluating Detail Losses and Additive Impairments,” in IEEE Transactions on Multimedia, vol. 13, No. 5, pp. 935-949, Oct. 2011, doi: 10.1109/TMM.2011.2152382., 15 pages. [cited by applicant]
Yang et al., “Perceptual Quality Assesment of Screen Content Images,” in IEEE Transactions on Image Processing, vol. 24, No. 11, pp. 4408-4421, Nov. 2015, doi: 10.1109/TIP.2015.2465145, 14 pages. [cited by applicant]