IP Library › Granted Patent US 12,206,862
Granted Patent B2
US 12,206,862 · App. 17/382,154 · Granted Jan 21, 2025

Methods for non-reference video-quality prediction

Inventor: Minhua Zhou (San Diego, CA)
Assignee: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
H04N19/154G06N3/044G06N3/063G06N5/04H04N19/164H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,206,862
App. No.
17/382,154
Granted
Jan 21, 2025
Kind
B2
Abstract

A system for non-reference video-quality prediction includes a video-processing block to receive an input bitstream and to generate a first vector, and a neural network to provide a predicted-quality vector after being trained using training data. The training data includes the first vector and a second vector, and elements of the first vector include high-level features extracted from a high-level syntax processing of the input bitstream.

Claims (33)

1. A system, comprising:

a video decoder configured to receive and decode an input bitstream to reconstruct a picture from the input bitstream, and to generate a first vector comprising features extracted from the input bitstream by the video decoder and features determined by the video decoder during reconstruction of the picture; and

a neural network configured to generate a second vector, based on the first vector, comprising one or more metrics representing a predicted quality of the picture reconstructed by the video decoder,

wherein one or more parameters of the neural network are trained by:

computing a first training vector based on a test bitstream that has been decoded,

generating, using the video decoder, a second training vector comprising features extracted from the test bitstream and determined by the video decoder,

generating, using the neural network, a third training vector based on the second training vector,

computing a prediction loss between the first training vector and the third training vector, and

updating the one or more parameters based on the prediction loss.

2. The system of claim 1 , wherein the video decoder is configured to extract the features from the input bitstream by parsing syntax elements in the input bitstream.

3. The system of claim 2 , wherein the syntax elements comprise at least one of sequence parameter sets, picture parameter sets, video parameter sets, picture headers, slice headers, adaptation parameter sets, or supplemental enhancement information messages.

4. The system of claim 2 , wherein the features extracted from the input bitstream comprise at least one of a transcode indicator, a codec type, a picture coding type, a picture resolution, a frame rate, a bit depth, a chroma format, a compressed picture size, a high-level quantization parameter, an average temporal distance, or a temporal layer identifier.

5. The system of claim 1 , wherein the features determined by the video decoder during reconstruction of the picture comprise at least one of a percentage of intra-coded blocks in the picture, a percentage of inter-coded blocks in the picture, an average block-level quantization parameter for the picture, a maximum block-level quantization parameter for the picture, a minimum block-level quantization for the picture, a standard deviation of horizontal-motion vectors for the picture, an average motion-vector size for the picture, an average absolute amplitude of low-frequency inverse quantized transform coefficients for the picture, an average absolute amplitude of high-frequency inverse quantized-transform coefficients for the picture, a standard deviation of a prediction residual for the picture, a root-mean-squared error value between reconstructed pictures before and after in-loop filters, a standard deviation of the reconstructed picture after in-loop filters, or an edge sharpness of the reconstructed picture after in-loop filters.

6. The system of claim 1 , wherein the one or more metrics representing the predicted quality of the reconstructed picture comprise at least one of a peak signal noise ratio, a structural similarity index measure, a multiscale similarity index measure, a video multimethod assessment fusion, or a mean opinion score.

7. The system of claim 6 , wherein an output layer of the neural network is configured to generate root-mean-squared-error values.

8. The system of claim 7 , wherein the root-mean-squared-error values are converted into peak signal noise ratio values.

9. A method, comprising:

receiving and decoding, by a video decoder, an input bitstream to reconstruct a picture from the input bitstream;

generating, by the video decoder, a first vector comprising features extracted from the input bitstream by the video decoder and determined during reconstruction of the picture by the video decoder;

generating, using a neural network, a second vector comprising one or more metrics representing a predicted quality of the reconstructed picture, wherein the second vector is based on the first vector,

wherein one or more parameters of the neural network are trained by:

computing a first training vector based on a test bitstream that has been decoded;

generating, using the video decoder, a second training vector comprising features extracted from the test bitstream and determined by the video decoder;

generating, using the neural network, a third training vector based on the second training vector;

computing a prediction loss between the first training vector and the third training vector; and

updating the one or more parameters based on the prediction loss.

10. The method of claim 9 , wherein the features are extracted from the input bitstream by parsing syntax elements in the input bitstream.

11. The method of claim 10 , wherein the syntax elements comprise at least one of sequence parameter sets, picture parameter sets, video parameter sets, picture headers, slice headers, adaptation parameter sets, or supplemental enhancement information messages.

12. The method of claim 10 , wherein the features extracted from the input bitstream comprise at least one of a transcode indicator, a codec type, a picture coding type, a picture resolution, a frame rate, a bit depth, a chroma format, a compressed picture size, a high-level quantization parameter, an average temporal distance, or a temporal layer identifier.

13. The method of claim 9 , wherein the features determined by the video decoder during reconstruction of the picture comprise at least one of a percentage of intra-coded blocks in the picture, a percentage of inter-coded blocks in the picture, an average block-level quantization parameter for the picture, a maximum block-level quantization parameter for the picture, a minimum block-level quantization for the picture, a standard deviation of horizontal-motion vectors for the picture, an average motion-vector size for the picture, an average absolute amplitude of low-frequency inverse quantized transform coefficients for the picture, an average absolute amplitude of high-frequency inverse quantized-transform coefficients for the picture, a standard deviation of a prediction residual for the picture, a root-mean-squared error value between reconstructed pictures before and after in-loop filters, a standard deviation of the reconstructed picture after in-loop filters, or an edge sharpness of the reconstructed picture after in-loop filters.

14. The method of claim 9 , wherein the one or more metrics representing the predicted quality of the reconstructed picture comprise at least one of a peak signal noise ratio, a structural similarity index measure, a multiscale similarity index measure, a video multimethod assessment fusion, or a mean opinion score.

15. The method of claim 14 , wherein an output layer of the neural network is configured to generate root-mean-squared-error values.

16. The method of claim 15 , wherein the root-mean-squared-error values are converted into peak signal noise ratio values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2023
From: ZHOU, MINHUA
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 062963/0636 →
Continuity (1)
Related Publication 20230024037A1 · Jan 26, 2023
References Cited (11)
US 11250546B2 · Pu · 2022 [cited by examiner]
US 11288770B2 · Kim · 2022 [cited by examiner]
US 20130293725A1 · Zhang · 2013 [cited by examiner]
US 20180084280A1 · Thiagarajan · 2018 [cited by examiner]
US 20190258902A1 · Colligan · 2019 [cited by examiner]
US 20210042882A1 · Kim · 2021 [cited by examiner]
US 20210385502A1 · Dinh · 2021 [cited by examiner]
Jiang et al., “No-Reference Perceptual Video Quality Measurement for High Definition Videos Based on an Artificial Neural Network,” 2008 International Conference on Computer and Electrical Engineering, Dec. 2008, pp. 42… [cited by applicant]
Shahid et al., “A reduced complexity no-reference artificial neural network based video quality predictor,” 4th International Congress on Image and Signal Processing, Oct. 2011, pp. 517-521. [cited by applicant]
Kang et al., “Convolutional Neural Networks for No-Reference Image Quality Assessment,” 2014 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2014, pp. 1733-1740. [cited by applicant]
Extended European Search Report from European Patent Application No. 22185436.7, dated Dec. 14, 2022, 10 pages. [cited by applicant]
Cited By (1)
US 12,694,660