IP Library Granted Patent US 12,501,084
Granted Patent B2
US 12,501,084 · App. 18/371,852 · Granted Dec 16, 2025

Efficient two-pass encoding scheme for adaptive live streaming

Inventors: Vignesh V. Menon (Klagenfurt am Wörthersee, AT); Hadi Amirpour (Klagenfurt am Wörthersee, AT); Christian Timmerer (Klagenfurt am Wörthersee, AT)
Assignee: Bitmovin GmbH
H04N21/2343H04N21/23418H04N21/2187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,501,084
App. No.
18/371,852
Granted
Dec 16, 2025
Kind
B2
Abstract

Techniques for efficient two-pass encoding for live streaming are described herein. A method for efficient two-pass encoding may include extracting low-complexity features of a video segment, predicting an optimized constant rate factor (CRF) for the video segment using the low-complexity features, and encoding the video segment with the optimized CRF at a target bitrate. A system for efficient two-pass encoding may include a feature extraction module configured to extract low-complexity features from a video segment, a neural network configured to predict an optimized CRF as a function of the low-complexity features and a target bitrate, and an encoder configured to encode the video segment using the optimized CRF at the target bitrate.

Claims (31)

1 . A method for efficient two-pass encoding comprising:

extracting low-complexity features of a video segment;

predicting, by a neural network, an optimized constant rate factor (CRF) for the video segment using the low-complexity features;

receiving, by the neural network, a bitrate set, a resolution set, and a framerate set, wherein predicting the optimized CRF is further based on a target bitrate from the bitrate set, a resolution from the resolution set, and a framerate from the framerate set; and

encoding the video segment with the optimized CRF at a target bitrate.

2 . The method of claim 1 , wherein the extracting low-complexity features of the video segment comprises computing for each of a plurality of resolutions and framerates one or both of a spatial complexity metric and a temporal complexity metric.

3 . The method of claim 1 , wherein the low-complexity features comprise at least a Discrete Cosine Transform (DCT)-energy-based spatial feature and a DCT-energy-based temporal feature.

4 . The method of claim 1 , wherein the extracting the low-complexity features of the video segment comprises a first-pass capped variable bitrate encoding.

5 . The method of claim 1 , wherein the neural network comprises a shallow network having as few as two hidden layers.

6 . The system for efficient two-pass encoding comprising:

a processor; and

a memory comprising program instructions executable by the processor to cause the processor to implement:

a feature extraction module configured to extract a low-complexity feature from a video segment;

a neural network configured to predict an optimized constant rate factor (CRF) as a function of the low-complexity feature of the video segment and a target bitrate; and

an encoder configured to encode the video segment using the optimized CRF at the target bitrate,

wherein the neural network is trained to predict the optimized CRF as a function of the target bitrate for each of a set of resolutions and each of a set of framerates.

7 . The system of claim 6 , wherein the low-complexity feature comprises at least a Discrete Cosine Transform (DCT)-energy-based spatial feature and a DCT-energy-based temporal feature.

8 . The system of claim 6 , wherein the low-complexity feature comprises one or both of a spatial complexity metric and a temporal complexity metric.

9 . The system of claim 6 , wherein the neural network is trained to determine a minimum CRF and a maximum CRF for a target codec and a target bitrate.

10 . The system of claim 6 , wherein the neural network comprises an input layer, two or more hidden layers, and an output layer.

11 . A distributed computing system comprising:

a distributed database configured to store a plurality of video segments, a neural network, a plurality of bitrate ladders, and a codec; and

one or more processors configured to:

extract low-complexity features of a video segment;

predict, by the neural network, an optimized constant rate factor (CRF) for the video segment using the low-complexity features; and

encode the video segment with the optimized CRF at a target bitrate,

wherein the neural network is trained to predict the optimized CRF as a function of the target bitrate for each of a set of resolutions and each of a set of framerates.

12 . The system of claim 11 , wherein the low-complexity feature comprises at least a Discrete Cosine Transform (DCT)-energy-based spatial feature and a DCT-energy-based temporal feature.

13 . The system of claim 11 , wherein the low-complexity feature comprises one or both of a spatial complexity metric and a temporal complexity metric.

14 . The system of claim 11 , wherein the neural network is trained to determine a minimum CRF and a maximum CRF for a target codec and a target bitrate.

15 . The system of claim 11 , wherein the neural network comprises an input layer, two or more hidden layers, and an output layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: MENON, VIGNESH V.; AMIRPOUR, HADI; TIMMERER, CHRISTIAN
To: BITMOVIN, GMBH
Reel/Frame 064999/0153 →
Continuity (2)
Provisional Application 63409922 · Sep 26, 2022
Related Publication 20240114183A1 · Apr 4, 2024
References Cited (43)
US 10104413B2 · Phillips et al. · 2018 [cited by applicant]
US 10419773B1 · Wei et al. · 2019 [cited by applicant]
US 10499081B1 · Wang et al. · 2019 [cited by applicant]
US 10798399B1 · Wei et al. · 2020 [cited by applicant]
US 10958947B1 · Wei et al. · 2021 [cited by applicant]
US 11445168B1 · Wei et al. · 2022 [cited by applicant]
US 20050018881A1 · Peker · 2005 [cited by examiner]
US 20100189183A1 · Gu et al. · 2010 [cited by applicant]
US 20110305273A1 · He et al. · 2011 [cited by applicant]
US 20120147958A1 · Ronca · 2012 [cited by applicant]
US 20130089142A1 · Begen et al. · 2013 [cited by applicant]
US 20130282917A1 · Reznik et al. · 2013 [cited by applicant]
US 20160073106A1 · Su · 2016 [cited by applicant]
US 20160134881A1 · Wang · 2016 [cited by applicant]
US 20170078574A1 · Puntambekar et al. · 2017 [cited by applicant]
US 20170078686A1 · Coward et al. · 2017 [cited by applicant]
US 20180014050A1 · Phillips et al. · 2018 [cited by applicant]
US 20180338146A1 · John · 2018 [cited by applicant]
US 20190028745A1 · Katsavounidis · 2019 [cited by applicant]
US 20190075301A1 · Chou et al. · 2019 [cited by applicant]
US 20190132591A1 · Zhang · 2019 [cited by examiner]
US 20190289296A1 · Kottke · 2019 [cited by examiner]
US 20200412784A1 · Yamagishi et al. · 2020 [cited by applicant]
US 20230012862A1 · Kossentini · 2023 [cited by examiner]
Bentaleb et al., “A Survey on Bitrate Adaptation Schemes for Streaming Media Over HTTP,”, IEEE Communications Surveys & Tutorials, vol. 21, No. 1, 2019, pp. 562-585. [cited by applicant]
Jain et al., “Throughput Fairness Index: An Explaination”, 1984, pp. 13. [cited by applicant]
Mehrabi et al., “Edge Computing Assisted Adaptive Mobile Video Streaming”, IEE Transactions on Mobile Computing, vol. 18, No. 4, Apr. 2019, pp. 787-800. [cited by applicant]
Lederer et al., “Dynamic Adaptive Streaming over HTTP Dataset”, Proceedings of the 3rd Multimedia Systems Conference, Feb. 2012, pp. 89-94. [cited by applicant]
Ericsson, “Ericsson Mobility Report”, Nov. 2019, pp. 1-36. [cited by applicant]
ETSI, “Mobile Edge Computing a Key Technology Towards 5G”, ETSI White Paper No. 11, Sep. 2015, pp. 1-16. [cited by applicant]
Nguyen et al., “Adaptation Method for Video Streaming over HTTP/2”, IEICE Communications Express Comex, vol. 1, pp. 1-6, https://www.researchgate.netpublication/292213198_Adaptation_Method_for_Video_Streaming_over_HTTP2… [cited by applicant]
3GPP “3GPP TS 26.247. Progressive Download and Dynamic Adaptive Streaming over HTTP (3GP-DASH)”, 2015, pp. 1, https://portal.3gpp.org/desktopmodules/Specifications/SpecificationDetails.aspx?specificationId=1444. [cited by applicant]
Gernot Zwantschko, “What is Per-Title Encoding? How to Efficiently Compress Video”, Bitmovin, pp. 1-14, https://bitmovin.com/per-title-encoding/. [cited by applicant]
V.V Menon et al., “Efficient Content-Adaptive Feature-Based Shot Detection for HTTP Adaptive Streaming” IEEE, May 20, 2021, pp. 1-2, https://www.youtube.com/watch?v=jkA1R0shpTc. [cited by applicant]
Liu et al., “Video Super-Resolution Based on Deep Learning: A Comprehensive Survey”, arXiv:2007.12928v3 [cs.CV], Mar. 16, 2022, pp. 1-33. [cited by applicant]
Jon Dahl, “Instant Per-Title Encoding”, Mux, Apr. 17, 2018, pp. 1-8, https://mux.com/blog/instant-per-title-encoding/. [cited by applicant]
Ledig et al., “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”, arXiv:1609.04802, May 25, 2017, pp. 1-19, http://arxiv.org/abs/1609.04802. [cited by applicant]
Mishra et al., “A Survey on Deep Neural Network Compression: Challenges, Overview, and Solutions”, arXiv:2010.03954, Oct. 5, 2020, pp. 1-19, https://arxiv.org/abs/2010.03954. [cited by applicant]
Li et al., “Toward a Practical Perceptual Video Quality Metric”, Netflix Technology Blog, Jun. 5, 2016, pp. 1-23, https://netflixtechblog.com/toward-a-practical-perceptual-video-quality-metric-653f208b9652. [cited by applicant]
Menon et al., “ETPS: Efficient Two-pass Encoding Scheme for Adaptive Live Streaming,” Athena, https://www.youtube.com/watch?v=-pb3VJtrBN4, Oct. 16-19, 2022, pp. 1-2. [cited by applicant]
Wiegand et al., “Overview of the H.264/AVC Video Coding Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Jul. 2003, pp. 560-576. [cited by applicant]
Zupancic et al., “Two-Pass Rate Control for Improved Quality of Experience in UHDTV Delivery,” IEEE Journal of Selected Topics in Signal Processing, 2016, pp. 1-13. [cited by applicant]
Wang et al., “SSIM-Motivated Two-Pass VBR Coding for HEVC,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 27, No. 10, Oct. 2017, pp. 2189-2203. [cited by applicant]