IP Library › Granted Patent US 12,739,399
Granted Patent B2
US 12,739,399 · App. 18/650,638 · Granted Sep 15, 2026

Quality-based processing of video

Inventors: Sam Tak Wu Kwong (Kowloon, HK); Yunhao Mao (Kowloon, HK); Shiqi Wang (Kowloon, HK)
Assignee: City University of Hong Kong
H04N19/146H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,399
App. No.
18/650,638
Granted
Sep 15, 2026
Kind
B2
Abstract

There is provided a computer-implemented method for processing a video. The computer-implemented method includes: (a) determining a target frame-level quality required for a frame of the video to be encoded, the determining of the target frame-level quality is based on, at least, a rate-quantization (R-Q) model that relates bit-rate and quantization step size and a quality-quantization model that relates quality measure and the quantization step size; and (b) determining one or more coding parameters for encoding the frame based on the determined target frame-level quality.

Claims (68)

1 . A computer-implemented method for processing a video, comprising:

(a) determining a target frame-level quality required for a frame of the video to be encoded, the determining of the target frame-level quality is based on, at least, a rate-quantization (R-Q) model that relates bit-rate and quantization step size and a quality-quantization model that relates quality measure and the quantization step size; and

(b) determining one or more coding parameters for encoding the frame based on the determined target frame-level quality,

wherein the R-Q model is defined by

R

=

γ

Q

,

where R is bit-rate, Q is quantization step size, and γ is model parameter of the R-Q model,

wherein the quality-quantization model comprises a DISTS-quantization (D-Q) model that relates DISTS value and the quantization step size, and the D-Q model is defined as D=αQ 62 , where D is DISTS value, Q is quantization step size, and α and β are model parameters of the D-Q model, and

wherein the model parameters of the R-Q model and the D-Q model are updated by actual coding results for encoding a next frame.

2 . The computer-implemented method of claim 1 , further comprises determining a target GOP-level quality required for a GOP of the video, the GOP comprising a plurality of frames including the frame to be encoded, and

wherein the determining of the target frame-level quality is further based on the determined target GOP-level quality.

3 . The computer-implemented method of claim 2 , wherein the determining of the target frame-level quality required for the frame of the video comprises distributing or allocating at least part of the target GOP-level quality to the plurality of frames of the GOP.

4 . The computer-implemented method of claim 2 , wherein the determining of the target frame-level quality required for the frame of the video comprises determining the target frame-level quality while optimizing a GOP-level rate-distortion (R-D) cost function.

5 . The computer-implemented method of claim 4 , wherein the GOP-level rate-distortion cost function is defined based on, at least, a GOP-level Lagrangian multiplier for the GOP.

6 . The computer-implemented method of claim 5 ,

wherein the GOP-level Lagrangian multiplier is related to the target GOP-level quality through the R-Q model and the D-Q model;

wherein the determining of the target frame-level quality required for the frame of the video comprises determining the target frame-level quality required for the frame of the video based on the GOP-level Lagrangian multiplier; or

wherein the GOP-level Lagrangian multiplier is related to the target GOP-level quality through the R-Q model and the D-Q model, and the determining of the target frame-level quality required for the frame of the video comprises determining the target frame-level quality required for the frame of the video based on the GOP-level Lagrangian multiplier.

7 . The computer-implemented method of claim 1 , wherein the one or more the coding parameters comprises a quantization parameter and a Lagrangian multiplier.

8 . The computer-implemented method of claim 7 , wherein the determining of the quantization parameter in (b) is based on

Q

=

(

D

α

)

1

β

⁢

and

⁢

QP

=

log

X

(

Q

)

×

A

+

B

where D is the target frame-level quality represented as a target frame-level DISTS value, Q is the quantization step size, α and β are model parameters of the D-Q model, QP is the quantization parameter, A, B, and X are constants.

9 . The computer-implemented method of claim 8 , wherein the determining of the Lagrangian multiplier in (b) is based on

λ

=

C

×

D

QP

E

where λ is the Lagrangian multiplier, QP is the quantization parameter, C, D, and E are constants.

10 . The computer-implemented method of claim 1 , further comprising:

(c) encoding the frame based on the one or more determined coding parameters.

11 . The computer-implemented method of claim 10 , wherein the encoding in (c) is performed based on versatile video coding (VVC) based technique.

12 . The computer-implemented method of claim 10 , further comprising:

(d) determining, based on the encoding of the frame, an output bit-rate and an output quality of the frame; and

(e) updating, based on the determined output bit-rate and output quality, the model parameters of the R-Q model and the quality-quantization model.

13 . The computer-implemented method of claim 12 , wherein the updating in (e) is performed based on a gradient descent update method.

14 . The computer-implemented method of claim 12 , further comprising:

performing or repeating steps (a) to (e) for multiple frames of the video.

15 . A system for processing a video, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of to the computer-implemented method of claim 1 .

16 . A non-transitory computer readable medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to execute the computer-implemented method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2024
From: KWONG, SAM TAK WU; MAO, YUNHAO; WANG, SHIQI
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 067269/0357 →
Continuity (2)
Provisional Application 63502444 · May 16, 2023
Related Publication 20240388718A1 · Nov 21, 2024
References Cited (31)
US 6831947B2 · Ribas Corbera · 2004 [cited by applicant]
US 8532169B2 · Wang et al. · 2013 [cited by applicant]
US 8588296B2 · Yang et al. · 2013 [cited by applicant]
US 9860543B2 · Xu et al. · 2018 [cited by applicant]
US 10542262B2 · Gao et al. · 2020 [cited by applicant]
US 10560696B2 · Gao et al. · 2020 [cited by applicant]
US 11025914B1 · Yuen et al. · 2021 [cited by applicant]
US 20030031128A1 · Kim · 2003 [cited by examiner]
US 20050175109A1 · Vetro · 2005 [cited by examiner]
US 20160301931A1 · Wen · 2016 [cited by examiner]
US 20210400273A1 · Rapaka · 2021 [cited by examiner]
CN 106416251 · 2020 [cited by applicant]
CN 107113432 · 2020 [cited by applicant]
JP 6019189 · 2016 [cited by applicant]
B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the versatile video coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, … [cited by applicant]
G. J. Sullivan, J. Ohm, W. Han, and T. Wiegand, “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668, 2012. [cited by applicant]
A. Wieckowski, J. Brandenburg, T. Hinz, C. Bartnik, V. George, Hege, C. Helmrich, A. Henkel, C. Lehmann, C. Stoffers, I. Zupancic, B. Bross, and D. Marpe, “Vvenc: An open and optimized Vvc encoder implementation,” in 20… [cited by applicant]
L. Li, B. Li, H. Li, and C. W. Chen, “A-domain optimal bit allocation algorithm for high efficiency video coding,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, No. 1, pp. 130-142, 2018. [cited by applicant]
M. Zhou, X. Wei, S. Kwong, W. Jia, and B. Fang, “Rate control method based on deep reinforcement learning for dynamic video sequences in HEVC,” IEEE Transactions on Multimedia, vol. 23, pp. 1106-1121, 2020. [cited by applicant]
Z. He, Y. Kim, and S. K. Mitra, “Low-delay rate control for DCT video coding via p-domain source modeling,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 11, No. 8, pp. 928-940, 2001. [cited by applicant]
B. Li, H. Li, L. Li, and J. Zhang, “λ-domain rate control algorithm for high efficiency video coding,” IEEE Transactions on Image Processing, vol. 23, No. 9, pp. 3841-3854, 2014. [cited by applicant]
Y. Mao, M. Wang, S. Wang, and S. Kwong, “High efficiency rate control for versatile video coding based on composite cauchy distribution,” IEEE Transactions on Circuits and Systems for Video Technology, 2021. [cited by applicant]
F. Liu and Z. Chen, “Multi-objective optimization of quality in vvc rate control for low-delay video coding,” IEEE Transactions on Image Processing, vol. 30, pp. 4706-4718, 2021. [cited by applicant]
M. Zhou, X. Wei, C. Ji, T. Xiang, and B. Fang, “Optimum quality control algorithm for versatile video coding,” IEEE Transactions on Broadcasting, vol. 68, No. 3, pp. 582-593, 2022. [cited by applicant]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing, vol. 13, No. 4, pp. 600-612, 2004. [cited by applicant]
K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assessment: unifying structure and texture similarity,” IEEE transactions on pattern analysis and machine intelligence, 2020. [cited by applicant]
H. Choi, J. Yoo, J. Nam, D. Sim, and I. V. Bajic, “Pixel-wise unified rate-quantization model for multi-level rate control,” IEEE Journal of Selected Topics in Signal Processing, vol. 7, No. 6, pp. 1112-1123, 2013. [cited by applicant]
S. Ma, W. Gao, and Y. Lu, “Rate-distortion analysis for 264/AVC video coding and its application to rate control,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, No. 12, pp. 1533-1544, 2005. [cited by applicant]
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural informatio… [cited by applicant]
M. Sugawara, S.-Y. Choi, and D. Wood, “ultra-high-definition television (Rec. ITU-R BT.2020): A generational leap in the evolution of television [standards in a nutshell],” IEEE Signal Processing Magazine, vol. 31, No. … [cited by applicant]
F. Bossen, J. Boyce, K. Suehring, X. Li, and V. Seregin, “JVET common test conditions and software reference configurations for SDR video,” JVET T2010, Oct. 2020. [cited by applicant]