IP Library Granted Patent US 12,355,939
Granted Patent B2
US 12,355,939 · App. 18/367,234 · Granted Jul 8, 2025

Content-adaptive encoder preset prediction for adaptive live streaming

Inventors: Vignesh V. Menon (Klagenfurt am Wörthersee, AT); Hadi Amirpour (Klagenfurt am Wörthersee, AT); Christian Timmerer (Klagenfurt am Wörthersee, AT)
Assignee: Bitmovin GmbH
H04N19/103H04N19/14H04N19/42H04N21/2187H04N21/23418H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,355,939
App. No.
18/367,234
Granted
Jul 8, 2025
Kind
B2
Abstract

Techniques for content-adaptive encoder preset prediction for adaptive live streaming are described herein. A method for content-adaptive encoder preset prediction for adaptive live streaming includes performing video complexity feature extraction on a video segment to extract complexity features such as an average texture energy, an average temporal energy, and an average lumiscence. These inputs may be provided to an encoding time prediction model, along with a bitrate ladder, a resolution set, a target video encoding speed, and a number of CPU threads for the video segment, to predict an encoding time, and an optimized encoding preset may be selected for the video segment by a preset selection function using the predicted encoding time. The video segment may be encoded according to the optimized encoding preset.

Claims (34)

1. A method for content-adaptive encoder preset prediction for adaptive live streaming, the method comprising:

performing video complexity feature extraction on a video segment, including extracting a complexity feature of the video segment, the complexity feature comprising one, or a combination, of an average texture energy (E), an average temporal energy (h), and an average lumiscence (L);

receiving by an encoding time prediction model a plurality of inputs comprising the complexity feature, a bitrate ladder, a resolution set, a target video encoding speed, and a number of CPU threads for the video segment;

predicting an encoding time by an encoding time prediction model using the plurality of inputs; and

selecting an optimized encoding preset for the video segment by a preset selection function using the encoding time.

2. The method of claim 1 , wherein the encoding time prediction model and the preset selection function comprise a convolutional neural network.

3. The method of claim 1 , wherein the bitrate ladder and resolution set are provided in logarithmic scale to the encoding time prediction model.

4. The method of claim 1 , wherein the encoding time prediction model is trained on a set of encoder presets for a target encoder, the set of encoder presets ranging from a minimum encoder preset to a maximum encoder preset.

5. The method of claim 1 , wherein the encoding time prediction model is trained using a gradient boosting framework.

6. The method of claim 1 , wherein selecting the optimized encoding preset comprises selecting a preset such that the encoding time is corresponds to the target video encoding speed.

7. The method of claim 1 , wherein the complexity feature comprises a low-complexity spatial feature and a low-complexity temporal feature.

8. The method of claim 1 , further comprising encoding a plurality of representations of the video segment using the optimized encoding preset.

9. The method of claim 8 , wherein the plurality of representations comprises a representation for each bitrate in the bitrate ladder.

10. A distributed computing system comprising:

a distributed database configured to store a plurality of video segments, a neural network, a plurality of bitrate ladders, and a codec; and

one or more processors configured to:

perform video complexity feature extraction on a video segment, including extracting a complexity feature of the video segment, the complexity feature comprising one, or a combination, of an average texture energy (E), an average temporal energy (h), and an average lumiscence (L);

receive by an encoding time prediction model a plurality of inputs comprising the complexity feature, a bitrate ladder, a resolution set, a target video encoding speed, and a number of CPU threads for the video segment;

predict an encoding time by an encoding time prediction model using the plurality of inputs; and

select an optimized encoding preset for the video segment by a preset selection function using the encoding time.

11. The distributed computing system of claim 10 , wherein the neural network is configured to implement the encoding time prediction model.

12. The distributed computing system of claim 10 , wherein the neural network is configured to implement the preset selection function.

13. The distributed computing system of claim 10 , wherein the optimized encoding preset is selected such that the encoding time corresponds to the target video encoding speed.

14. The distributed computing system of claim 10 , wherein the one or more processors are further configured to encode a plurality of representations of the video segment using the optimized encoding preset, each of the plurality of representations corresponding to a bitrate in the bitrate ladder.

15. A system for content-adaptive encoder preset prediction for adaptive live streaming, the system comprising:

a processor; and

a memory comprising program instructions executable by the processor to cause the processor to implement:

an encoding time prediction model configured to predict an encoding time for a video segment using a plurality of inputs comprising the complexity feature, a bitrate ladder, a resolution set, a target video encoding speed, and a number of CPU threads for the video segment;

a preset selection function configured to select an optimized encoding preset for the video segment using the encoding time; and

an encoder configured to encode the video segment using the optimized encoding preset.

16. The system of claim 15 , wherein the encoding time prediction model is trained on a set of encoder presets for a target encoder, the set of encoder presets ranging from a minimum encoder preset to a maximum encoder preset.

17. The system of claim 15 , wherein the encoding time prediction model is trained using a gradient boosting framework.

18. The system of claim 15 , wherein the encoder comprises a plurality of encoders, each of the plurality of encoders configured to encode the video segment to generate a representation corresponding to a bitrate in the bitrate ladder.

19. The system of claim 15 , wherein the preset selection function is configured to select the optimized encoding preset such that the encoding time to correspond with the target video encoding speed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: MENON, VIGNESH V.; AMIRPOUR, HADI; TIMMERER, CHRISTIAN
To: BITMOVIN, GMBH
Reel/Frame 064878/0643 →
Continuity (2)
Provisional Application 63406136 · Sep 13, 2022
Related Publication 20240098247A1 · Mar 21, 2024
References Cited (39)
US 10104413B2 · Phillips et al. · 2018 [cited by applicant]
US 10499081B1 · Wang et al. · 2019 [cited by applicant]
US 20100189183A1 · Gu et al. · 2010 [cited by applicant]
US 20110305273A1 · He et al. · 2011 [cited by applicant]
US 20120147958A1 · Ronca · 2012 [cited by applicant]
US 20130089142A1 · Begen et al. · 2013 [cited by applicant]
US 20130282917A1 · Reznik et al. · 2013 [cited by applicant]
US 20160073106A1 · Su · 2016 [cited by applicant]
US 20160134881A1 · Wang · 2016 [cited by applicant]
US 20170078686A1 · Coward et al. · 2017 [cited by applicant]
US 20180014050A1 · Phillips et al. · 2018 [cited by applicant]
US 20180338146A1 · John · 2018 [cited by applicant]
US 20190028745A1 · Katsavounidis · 2019 [cited by applicant]
US 20190075301A1 · Chou et al. · 2019 [cited by applicant]
US 20200412784A1 · Yamagishi et al. · 2020 [cited by applicant]
US 20220408097A1 · Lin · 2022 [cited by examiner]
Bentaleb et al., “A Survey on Bitrate Adaptation Schemes for Streaming Media Over HTTP,”, IEEE Communications Surveys & Tutorials, vol. 21, No. 1, 2019, pp. 562-585. [cited by applicant]
Jain et al., “Throughput Fairness Index: An Explaination”, 1984, pp.—13. [cited by applicant]
Mehrabi et al., “Edge Computing Assisted Adaptive Mobile Video Streaming”, IEE Transactions on Mobile Computing, vol. 18, No. 4, Apr. 2019, pp. 787-800. [cited by applicant]
Lederer et al., “Dynamic Adaptive Streaming over HTTP Dataset”, Proceedings of the 3rd Multimedia Systems Conference, Feb. 2012, pp. 89-94. [cited by applicant]
Ericsson, “Ericsson Mobility Report”, Nov. 2019, pp. 1-36. [cited by applicant]
ETSI, “Mobile Edge Computing A Key Technology Towards 5G”, ETSI White Paper No. 11, Sep. 2015, pp. 1-16. [cited by applicant]
Nguyen et al., “Adaptation Method for Video Streaming over HTTP/2”, IEICE Communications Express Comex, vol. 1, pp. 1-6, https://www.researchgate.net publication/292213198_Adaptation_Method_for_Video_Streaming_over_HTTP… [cited by applicant]
3GPP “3GPP TS 26.247. Progressive Download and Dynamic Adaptive Streaming over HTTP (3GP-DASH)”, 2015, pp. 1, https://portal.3gpp.org/desktopmodules/Specifications/SpecificationDetails.aspx?specificationId=1444. [cited by applicant]
Gernot Zwantschko, “What is Per-Title Encoding? How to Efficiently Compress Video”, BITMOVIN, pp. 1-14, https://bitmovin.com/per-title-encoding/. [cited by applicant]
V.V Menon et al., “Efficient Content-Adaptive Feature-Based Shot Detection for HTTP Adaptive Streaming” IEEE, May 20, 2021, pp. 1-2, https://www.youtube.com/watch?v=jkA1R0shpTc. [cited by applicant]
Liu et al., “Video Super-Resolution Based on Deep Learning: A Comprehensive Survey”, arXiv:2007.12928v3 [cs.CV], Mar. 16, 2022, pp. 1-33. [cited by applicant]
Jon Dahl, “Instant Per-Title Encoding”, MUX, Apr. 17, 2018, pp. 1-8, https://mux.com/blog/instant-per-title-encoding/. [cited by applicant]
Ledig et al., “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”, arXiv:1609.04802, May 25, 2017, pp. 1-19, http://arxiv.org/abs/1609.04802. [cited by applicant]
Mishra et al., “A Survey on Deep Neural Network Compression: Challenges, Overview, and Solutions”, arXiv:2010.03954, Oct. 5, 2020, pp. 1-19, https://arxiv.org/abs/2010.03954. [cited by applicant]
Li et al., “Toward A Practical Perceptual Video Quality Metric”, Netflix Technology Blog, Jun. 5, 2016, pp. 1-23, https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652. [cited by applicant]
Menon et al., “ETPS: Efficient Two-pass Encoding Scheme for Adaptive Live Streaming,” Athena, https://www.youtube.com/watch?v=-pb3VJtrBN4, Oct. 16-19, 2022, pp. 1-2. [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, pp. 1649-1668. [cited by applicant]
Chen et al., “XGBoost: A Scalable Tree Boosting System,” Machine Learing, Mar. 2016, pp. 785-794. [cited by applicant]
Menon et al., “Content-adaptive Encoder Preset Prediction for Adaptive Live Streaming,” arXiv:2210.10330, Oct. 2022, pp. 1-5. [cited by applicant]
Huangyuan et al., “Performance Evaluation of H.265/MPEG-HEVC Encoders for 4K Video Sequences,” APSIPA, 2014, pp. 1-8. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” arXiv:1502.03167, Mar. 2015, pp. 1-11. [cited by applicant]
Zvezdakov et al., “Machine-Learning-Based Method for Content-Adaptive Video Encoding,” IEEE Xplore, 2021, pp. 1-5. [cited by applicant]
Intel Corporation, “Accelerating x265 with Intel Advanced Vector Extensions 512” White Paper, Mar. 2018, pp. [cited by applicant]