IP Library › Granted Patent US 12,501,061
Granted Patent B2
US 12,501,061 · App. 18/440,013 · Granted Dec 16, 2025

Multivariate rate control for transcoding video content

Inventors: Sam John (Dublin, CA); Balineedu Adsumilli (Sunnyvale, CA); Akshay Gadde (Fremont, CA)
Assignee: GOOGLE LLC
H04N19/40H04N19/119H04N19/147H04N19/184H04N19/192
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,501,061
App. No.
18/440,013
Granted
Dec 16, 2025
Kind
B2
Abstract

A learning model is trained for rate-distortion behavior prediction against a corpus of a video hosting platform and used to determine optimal bitrate allocations for video data given video content complexity across the corpus of the video hosting platform. Complexity features of the video data are processed using the learning model to determine a rate-distortion cluster prediction for the video data, and transcoding parameters for transcoding the video data are selected based on that prediction. The rate-distortion clusters are modeled during the training of the learning model, such as based on rate-distortion curves of video data of the corpus of the video hosting platform and based on classifications of such video data. This approach minimizes total corpus egress and/or storage while further maintaining uniformity in the delivered quality of videos by the video hosting platform.

Claims (51)

1 . A method, comprising:

determining that a rate-distortion behavior of a video chunk of an input video stream uploaded to a video hosting platform is similar to a rate-distortion behavior of videos used to produce a rate-distortion cluster of a plurality of rate-distortion clusters, each corresponding to a different rate-distortion classification of videos of a corpus of the video hosting platform, based on a rate-distortion classification of the video chunk and a rate-distortion classification of video content to which the rate-distortion cluster corresponds;

selecting the rate-distortion cluster from amongst the plurality of rate-distortion clusters based on the determining that the rate-distortion behavior of the video chunk is similar to the rate-distortion behavior of the videos used to produce the rate-distortion cluster; and

transcoding the video chunk according to transcoding parameters selected based on one or more operating points of the rate-distortion cluster.

2 . The method of claim 1 , comprising:

selecting the transcoding parameters based on operating points of a centroid curve of the rate-distortion cluster.

3 . The method of claim 2 , wherein each operating point represents a bitrate available for transcoding and a quality resulting from using the bitrate, and wherein selecting the transcoding parameters based on the operating points of the centroid curve of the rate-distortion cluster comprises:

identifying, as an optimal operating point, one of the operating points of the centroid curve; and

selecting, as the transcoding parameters, parameters corresponding to the optimal operating point.

4 . The method of claim 2 , comprising:

determining centroid curves each for a different rate-distortion cluster of the plurality of rate-distortion clusters.

5 . The method of claim 1 , comprising:

determining the rate-distortion classification of the video chunk based on one or more complexity features of the video chunk.

6 . The method of claim 5 , comprising:

determining the one or more complexity features based on a pass log of an encoder used for encoding the input video stream.

7 . The method of claim 6 , wherein the pass log is received after a first pass encoding by the encoder, the method further comprising:

verifying, before a second pass encoding by the encoder, the selection of the transcoding parameters according to one or more transcoder constraints.

8 . The method of claim 5 , wherein determining that the rate-distortion behavior of the video chunk is similar to the rate-distortion behavior of the videos used to produce the rate-distortion cluster comprises:

determining, using a learning model trained to predict rate-distortion behaviors of video data at least some of the videos of the corpus of the video hosting platform, a correspondence of the video chunk to the rate-distortion cluster based on the one or more complexity features.

9 . The method of claim 8 , comprising:

training the learning model using the at least some of the videos.

10 . The method of claim 9 , wherein training the learning model using the at least some of the videos comprises:

determining rate-distortion curves for at least some of the videos; and

producing the plurality of rate-distortion clusters by clustering the rate-distortion curves based on similarities of complexity features of the at least some of the videos.

11 . A method, comprising:

receiving an input video stream uploaded to a video hosting platform;

determining that a rate-distortion behavior of a video chunk of the input video stream is similar to a rate-distortion behavior associated with a rate-distortion cluster of a plurality of rate-distortion clusters based on a rate-distortion classification of the video chunk and a rate-distortion classification of video content to which the rate-distortion cluster corresponds, wherein each rate-distortion cluster of the plurality of rate-distortion clusters corresponds to a different rate-distortion classification of videos of a corpus of the video hosting platform;

selecting the rate-distortion cluster from amongst the plurality of rate-distortion clusters based on the determining that the rate-distortion behavior of the video chunk is similar to the rate-distortion behavior associated with the rate-distortion cluster; and

transcoding the video chunk according to transcoding parameters selected based on one or more operating points of the rate-distortion cluster.

12 . The method of claim 11 , comprising:

selecting the transcoding parameters based on an optimal operating point of a centroid curve of the rate-distortion cluster.

13 . The method of claim 11 , comprising:

determining the rate-distortion classification of the video chunk.

14 . The method of claim 11 , comprising:

training a learning model using at least some of the videos of the corpus of the video hosting platform, wherein the learning model is used to determine that the rate-distortion behavior of the video chunk is similar to the rate-distortion behavior associated with the rate-distortion cluster.

15 . A method, comprising:

determining a similarity between a rate-distortion behavior of a video chunk of an input video stream and a rate-distortion behavior associated with a rate-distortion cluster of a plurality of rate-distortion clusters, each corresponding to a different rate-distortion classification of videos of a corpus of a video hosting platform, based on a rate-distortion classification of the video chunk and a rate-distortion classification of video content to which the rate-distortion cluster corresponds;

selecting the rate-distortion cluster from amongst the plurality of rate-distortion clusters based on the determining of the similarity between the rate-distortion behavior of the video chunk and the rate-distortion behavior associated with the rate-distortion cluster;

selecting transcoding parameters based on one or more operating points of the rate-distortion cluster; and

transcoding the video chunk according to the transcoding parameters.

16 . The method of claim 15 , comprising:

identifying the plurality of rate-distortion clusters using a learning model trained based on the corpus of the video hosting platform.

17 . The method of claim 16 , wherein identifying the plurality of rate-distortion clusters using the learning model trained based on the corpus of the video hosting platform comprises:

receiving a training data set including training video data from at least some of the videos of the corpus of the video hosting platform;

determining rate-distortion curves for the training video data; and

producing rate-distortion clusters by clustering the rate-distortion curves based on similarities of complexity features of the training video data.

18 . The method of claim 17 , comprising:

determining a centroid curve for each of the rate-distortion clusters.

19 . The method of claim 18 , wherein the centroid curve determined for each of the rate-distortion clusters includes a number of operating points, wherein each operating point represents a bitrate available for transcoding and a quality resulting from using the bitrate.

20 . The method of claim 18 , wherein selecting the transcoding parameters based on the rate-distortion cluster comprises:

selecting, as the transcoding parameters, parameters corresponding to an operating point of the centroid curve of the rate-distortion cluster.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2024
From: JOHN, SAM; ADSUMILLI, BALINEEDU; GADDE, AKSHAY
To: GOOGLE LLC
Reel/Frame 066449/0816 →
Continuity (2)
Continuation 17908352
Related Publication 20240187618A1 · Jun 6, 2024
References Cited (28)
US 8897370B1 · Wang · 2014 [cited by examiner]
US 10419773B1 · Wei et al. · 2019 [cited by applicant]
US 10623775B1 · Theis et al. · 2020 [cited by applicant]
US 11259040B1 · Sabui et al. · 2022 [cited by applicant]
US 20060222078A1 · Raveendran · 2006 [cited by applicant]
US 20120275511A1 · Shemer et al. · 2012 [cited by applicant]
US 20150063436A1 · Lasserre et al. · 2015 [cited by applicant]
US 20170078676A1 · Coward et al. · 2017 [cited by applicant]
US 20180109799A1 · De Cock et al. · 2018 [cited by applicant]
US 20200090069A1 · Mandt · 2020 [cited by examiner]
US 20200128242A1 · Aristarkhov · 2020 [cited by examiner]
US 20200273040A1 · Novick et al. · 2020 [cited by applicant]
US 20210049757A1 · Zhu et al. · 2021 [cited by applicant]
CN 110446048A · 2019 [cited by examiner]
EP 1677252A1 · 2006 [cited by examiner]
WO WO2006099082A2 · 2006 [cited by examiner]
WO WO2007028515A2 · 2007 [cited by examiner]
Patrick Le Callet, “On Perceptual Coding: Quality, Content Features and Complexity”, AOMedia Symposium 2019, 47 pgs. [cited by applicant]
Y. Wang, S. Inguva, and B. Adsumili, “YouTube UGC Dataset for Video Compression Research,” in IEEE International Workshop on Multimedia Signal Processing, 2019, 5 pgs. [cited by applicant]
M. Seufert, S. Egger, M. Slanina, T. Zinnder, T. Hoßfeld, and P. Tran-Gia, “A Survey on Quality of Experience of HTTP Adaptive Streaming,” IEEE Communications Surveys Tutorials, 2015, pp. 469-492. [cited by applicant]
A. Ortega and K. Ramchandran, “Rate-Distortion Methods for Image and Video Compression,” IEEE Signal Processing Magazine, Nov. 1998, pp. 23-50. [cited by applicant]
L. Toni, R. Aparicio-Pardo, K. Pires, G. Simon, A. Blanc, and P. Frossard, “Optimal Selection of Adaptive Streaming Representations,” ACM Trans. Multimedia Comput. Commun. Appl., vol. 11, No. 2s, Article 43, Feb. 2015, … [cited by applicant]
C. Chen, Y. Lin S. Benting, and A. Kokaram, “Optimized Transcoding For Large Scale Adaptive Streaming Using Playback Statistics,” IEEE International Conference on Image Processing, Oct. 2018, pp. 3269-3273. [cited by applicant]
Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi, et al., “An overview of core coding tools in the AV1 video codec,” in IEEE Picture Coding Symposium, 2018. [cited by applicant]
S. Ling, Y. Baveye, P. Le Callet, J. Skinner, and I. Katsavounidis, “Characterization of User Generated Content for Perceptually-Optimized Video Compression: Challenges, Observations and Perspectives,” Human Vision and … [cited by applicant]
C.-C. Chang and C.-J. Lin, “LIBSVM: A Library for Support Vector Machines,” ACM Transactions on Intelligent Systems and Technology, vol. 2, No. 3, Article 27, Apr. 2011, 27 pgs. [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/US2020/03354 dated Dec. 15, 2020, 14 pgs. [cited by applicant]
YouTube Engineering and Developers Blog: Making high quality video efficient, Apr. 24, 2018, 8 pgs. [cited by applicant]