IP Library Granted Patent US 11,076,153
Granted Patent B2
US 11,076,153 · App. 15/747,982 · Granted Jul 27, 2021

System and methods for joint and adaptive control of rate, quality, and computational complexity for video coding and video delivery

Inventors: Marios Stephanou Pattichis (Albuquerque, NM); Yuebing Jiang (Santa Clara, CA); Cong Zong (Albuquerque, NM); Gangadharan Esakki (Albuquerque, NM); Venkatesh Jatla (Albuquerque, NM); Andreas Panayides (Strovolos, CY)
H04N19/127H04N19/119H04N19/126H04N19/147H04N19/149H04N19/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,076,153
App. No.
15/747,982
Granted
Jul 27, 2021
Kind
B2
Abstract

System and methods for the joint control of reconstructed video quality, computational complexity and compression rate for intra-mode and inter-mode video encoding in HEVC. The invention provides effective methods for (i) generating a Pareto front for intra-coding by varying CTU parameters and the QP, (ii) generating a Pareto front for inter-coding by varying GOP configurations and the QP, (iii) real-time and offline Pareto model front estimation using regression methods, (iv) determining the optimal encoding configurations based on the Pareto model by root finding and local search, and (v) robust adaptation of the constraints and model updates at both the CTU and GOP levels.

Claims (155)

1. A method for real-time adaptive encoding digital video signals comprising:

(a) receiving an input video comprising a plurality of video segments;

(b) applying, to a video segment, real-time input constraints on: (1) video quality remaining above a minimum value Q, (2) bandwidth with bitrate remaining below a maximum value representing available bitrate, and (3) encoding frame rate with a number of frames per second (FPS) remaining above a minimum encoding rate value, to select initial candidate encoding configurations, wherein applying further comprises using pre-computed forward regression models, wherein the pre-computed forward regression models can vary based on an encoding scheme, and are given by:

log( Q )= a 0 +b 0 ·QP+ c 0 ·QP 2 ,

log(Bitrate)= a 1 +b 1 ·QP+ c 1 ·QP 2 ,

log(FPS)= a 2 +b 2 ·QP+ c 2 ·QP 2 ,  Equation (9.1)

wherein QP is a quantization parameter and a 0 , b 0 , c 0 , a 1 , b 1 , c 1 , a 2 , b 2 , c 2 represent regression coefficients determined using a training process that uses video segments similar to the video segments of the plurality;

(c) using the pre-computed forward regression models to derive inverse models to determine final candidate encoding configurations from the initial candidate encoding configurations;

(d) selecting an optimal encoding configuration from the final candidate encoding configurations, wherein the optimal encoding configuration satisfies constraints and achieves a maximum video quality, a minimum bandwidth, or a maximum frame rate, wherein the optimal encoding configuration comprises of a Group of Pictures (GOP) configuration and a Coding Tree Unit (CTU) configuration;

(e) encoding the video segment using the optimal encoding configuration; and

(f) repeating (b)-(e) for all video segments of the plurality of video segments.

2. The method of claim 1 further comprising creating off-line the pre-computed forward regression models.

3. The method of claim 2 , wherein creating further comprises:

inputting a plurality of videos that is composed of video segments;

encoding each video segment using different video encoding parameters, Coding Tree Unit configurations, and GOP configurations;

evaluating the video quality, required bitrate, and video encoding rate in frames per second for each video segment; and

learning the forward regression models that map the video encoding parameters, Coding Tree Unit configurations, and GOP configurations to the video quality, required bitrate, and video encoding rate over a training set of video segments.

4. The method of claim 1 , wherein the inverse models use Newton's algorithm to determine final candidate encoding configurations from the forward regression models and constraints on video quality, maximum bitrate, and minimum video encoding rate.

5. The method of claim 1 , wherein the optimal encoding configuration is one selected from the group: a maximum video encoding performance mode, a minimum bitrate mode, and a maximum video quality mode.

6. The method of claim 5 , wherein the maximum performance mode is defined according to:

min

c

C

T

subject

to

:

(

Q

Q

min

)

&

(

R

R

max

)

with C representing a set of video encoding configurations, R representing a number of bits per pixel, T representing encoding time per frame, and Q representing a measure of video quality.

7. The method of claim 5 , wherein the minimum bitrate mode is defined according to:

min

c

C

R

subject

to

:

(

Q

Q

min

)

&

(

T

T

max

)

with C representing a set of video encoding configurations, R representing a number of bits per pixel, T representing encoding time per frame, and Q representing a measure of video quality.

8. The method of claim 5 , wherein the maximum quality mode is defined according to:

min

c

C

Q

subject

to

:

(

T

T

max

)

&

(

R

R

max

)

with C representing a set of video encoding configurations, R representing a number of bits per pixel, T representing encoding time per frame, and Q representing a measure of video quality.

9. The method of claim 1 , wherein the forward regression model is defined in terms of a quantization parameter (QP), the GOP configuration, and the Coding Tree Unit configuration.

10. The method of claim 1 , wherein video constraints and the optimization modes are applied individually in a CTU or a GOP while staying within a budget.

11. The method of claim 10 , wherein the budget comprises a target bitrate (R target ) of a number of bits per second for each video frame according to the equation:

R target =N pixels ·bbp target

wherein N pixels is a number of pixels in each frame and bbp target is a required number of bits per pixel.

12. The method of claim 10 , wherein the budget comprises a target frame rate (T target ) of a total amount of time allocated to an entire frame according to the equation:

T target =N pixels ·time_per_pixel target

wherein N pixels is a number of pixels in each frame and time_per_pixel target is an encoding time per pixel.

13. The method of claim 10 , wherein the budget comprises a target video quality (Q target ) of a sum of squared error (SSE) for an entire frame according to the equation:

Q

target

=

2

2

·

bitDepth

·

N

p

i

x

e

l

s

1

0

P

S

N

R

/

1

0

wherein bitDepth is a number of bits used to represent each pixel, N pixels is a number of pixels in each frame and PSNR is Peak Signal-to-Noise Ratio.

14. The method of claim 1 , wherein (a)-(f) are applied to different video segments in a video delivery system, wherein the video delivery system can support live and on demand settings.

15. The method of claim 14 , wherein the video delivery system includes adaptive HTTP streaming (e.g., MPEG-DASH protocol) and RTP protocol based systems.

16. The method of claim 1 , wherein the encoding scheme comprises a GOP configuration.

17. The method of claim 1 , wherein the encoding scheme comprises encoding parameters that do not include QP.

18. The method of claim 1 , wherein the video quality is one selected from the group: structural similarity index measure (SSIM), peak signal-to-noise ratio (PSNR), and video multimethod assessment fusion (VMAF).

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2024
From: PATTICHIS, MARIOS STEPHANOU; JIANG, YUEBING; ZONG, CONG; ESAKKI, GANGADHARAN; PANAYIDES, ANDREAS; JATLA, VENKATESH
To: THE REGENTS OF THE UNIVERSITY OF NEW MEXICO
Reel/Frame 066982/0984 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2024
From: THE REGENTS OF THE UNIVERSITY OF NEW MEXICO
To: UNM RAINFOREST INNOVATIONS
Reel/Frame 066983/0395 →
CONFIRMATORY LICENSE Recorded Jun 17, 2019
From: UNIVERSITY OF NEW MEXICO
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 049493/0903 →
Continuity (2)
Provisional Application 62199438 · Jul 31, 2015
Related Publication 20180220133A1 · Aug 2, 2018
Cited By (4)
US 12,542,955 US 12,581,172 US 12,587,698 US 12,593,081