IP Library › Granted Patent US 12,262,032
Granted Patent B2
US 12,262,032 · App. 18/013,240 · Granted Mar 25, 2025

Reinforcement learning based rate control

Inventors: Jiahao Li (Beijing, CN); Bin Li (Beijing, CN); Yan Lu (Beijing, CN); Tom W. Holcomb (Sammamish, WA); Mei-Hsuan Lu (Taipei, TW); Andrey Mezentsev (Redmond, WA); Ming-Chieh Lee (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
H04N19/196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,262,032
App. No.
18/013,240
Granted
Mar 25, 2025
Kind
B2
Abstract

Implementations of the subject matter described herein provide a solution for rate control based on reinforcement learning. In this solution, an encoding state of a video encoder is determined, the encoding state being associated with encoding of a first video unit by the video encoder. An encoding parameter associated with rate control in the video encoder is determined by a reinforcement learning model and based on the encoding state of the video encoder. A second video unit different from the first video unit is encoded based on the encoding parameter. In this way, it is possible to achieve a better quality of experience (QOE) for real time communication with computation overhead being reduced.

Claims (66)

1. A computer-implemented method, comprising:

determining an encoding state of a video encoder, the encoding state associated with encoding a first video unit by the video encoder;

determining, by a reinforcement learning model and based on the encoding state of the video encoder, an encoding parameter associated with rate control for the video encoder;

encoding a second video unit different from the first video unit based on the encoding parameter; and

training the reinforcement learning model according to a reward for the encoding parameter based on the encoding of the second video unit, the reward being configured to penalize buffer overshooting and to increase as the encoding parameter results in a higher visual quality, wherein the reward is based on a base reward that:

has a negative value if buffer overshooting occurs;

increases as the encoding parameter decreases if buffer overshooting does not occur; and

is scaled by a scaling factor to obtain the reward.

2. The method of claim 1 , wherein determining the encoding parameter comprises:

determining, by the reinforcement learning model, an action based on the encoding state of the video encoder; and

mapping the action to the encoding parameter.

3. The method of claim 1 , wherein the encoding state associated with the encoding the first video unit comprises:

a state representing an outcome for encoding at least the first video unit;

a state of a buffer configured to buffer video units encoded by the video encoder before transmission; and

a state associated with a status of a network for transmitting the encoded video units.

4. The method of claim 3 , wherein the outcome for encoding at least the first video unit comprises the encoding parameter from the encoding of the first video unit and a size of the encoded first video unit, the state of the buffer comprises the usage of the buffer, and the state associated with the status of the network comprises a target bits per pixel.

5. The method of claim 4 , wherein the usage of the buffer comprises at least one of:

a ratio of an occupied space to maximum space of the buffer; and

remaining space of the buffer measured in video units.

6. The method of claim 1 , wherein the scaling factor is based on a ratio of a bandwidth associated with the encoding the second video unit to maximum bandwidth of a transmission channel.

7. The method of claim 1 , wherein the reinforcement learning model is trained by:

determining an action associated with the encoding parameter based on the encoding state of the video encoder;

determining an evaluation value for encoding state for the encoding the second video unit;

determining a value loss based on the reward and the evaluation value;

determining a policy loss based on the action; and

updating the reinforcement learning model based on the value loss and the policy loss.

8. The method of claim 1 , wherein the reinforcement learning model comprises a neural network of an agent, and wherein the neural network comprises:

at least one input fully connected layer configured to extract features from the encoding state;

at least one recurrent neural network coupled to receive the extracted features; and

at least one output fully connected layer configured to decide an action for the agent.

9. The method of claim 8 , wherein the neural network is trained based on an actor-critic architecture, the actor being configured to generate the action based on the encoding state, and the critic being configured to generate an evaluation value for the encoding state; and

wherein the actor and the critic share a common portion of the neural network comprising the at least one input fully connected layer and the at least one recurrent neural network.

10. The method of claim 1 , wherein the encoding parameter comprises at least one of a quantization parameter and a lambda parameter.

11. The method of claim 1 , wherein the video encoder is configured to encode screen content for real-time communication.

12. A device comprising:

a processor; and

a memory having instructions stored thereon for execution by the processor, the instructions for causing, when executed by the processor, the device to perform acts including:

determining an encoding state of a video encoder, the encoding state associated with encoding a first video unit by the video encoder;

determining, by a reinforcement learning model and based on the encoding state of the video encoder, an encoding parameter associated with rate control in the video encoder, wherein the reinforcement learning model has been trained based on a reward for the encoding parameter, the reward being configured to penalize buffer overshooting and to increase as the encoding parameter results in a higher visual quality, and wherein the reward is based on a base reward that:

has a negative value if buffer overshooting occurs;

increases as the encoding parameter decreases if the buffer overshooting does not occur; and

is scaled by a scaling factor to obtain the reward; and

encoding a second video unit different from the first video unit based on the encoding parameter.

13. The device of claim 12 , wherein the encoding state associated with the encoding the first video unit comprises:

a state representing an outcome for encoding at least the first video unit;

a state of a buffer configured to buffer video units encoded by the video encoder before transmission; and

a state associated with a status of a network for transmitting the encoded video units.

14. The device of claim 13 , wherein the outcome for encoding at least the first video unit comprises the encoding parameter from the encoding of the first video unit and a size of the encoded first video unit, the state of the buffer comprises the usage of the buffer, and the state associated with the status of the network comprises a target bits per pixel.

15. The device of claim 12 , wherein the scaling factor is based on a ratio of a bandwidth associated with encoding one or more video units to maximum bandwidth of a transmission channel.

16. The device of claim 14 , wherein the reinforcement learning model comprises a neural network of an agent, and wherein the neural network comprises:

at least one input fully connected layer configured to extract features from the encoding state;

at least one recurrent neural network coupled to receive the extracted features; and

at least one output fully connected layer configure to decide an action for the agent.

17. The device of claim 12 , wherein the video encoder is configured to encode screen content for real-time communication.

18. A computer-readable storage medium having program instructions stored thereon, the program instructions being executable by a processor to cause the processor to perform acts comprising:

determining an encoding state of a video encoder, the encoding state associated with encoding a first video unit by the video encoder,

determining, by a reinforcement learning model and based on the encoding state of the video encoder, an encoding parameter associated with rate control in the video encoder, wherein the reinforcement learning model has been trained based on a reward for the encoding parameter, the reward being configured to penalize buffer overshooting and to increase as the encoding parameter results in a higher visual quality, and wherein the reward is based on a base reward that:

has a negative value if buffer overshooting occurs;

increases as the encoding parameter decreases if the buffer overshooting does not occur; and

is scaled by a scaling factor to obtain the reward; and

encoding a second video unit different from the first video unit based on the encoding parameter.

19. The computer-readable storage medium of claim 18 , wherein the encoding state associated with the encoding the first video unit comprises:

a state representing an outcome for encoding at least the first video unit;

a state of a buffer configured to buffer video units encoded by the video encoder before transmission; and

a state associated with a status of a network for transmitting the encoded video units.

20. The computer-readable storage medium of claim 19 , wherein the outcome for encoding at least the first video unit comprises the encoding parameter from the encoding of the first video unit and a size of the encoded first video unit, the state of the buffer comprises the usage of the buffer, and the state associated with the status of the network comprises a target bits per pixel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2023
From: LI, JIAHAO; LI, BIN; LU, YAN; HOLCOMB, TOM W.; LU, MEI-HSUAN; MEZENTSEV, ANDREY; LEE, MING-CHIEH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062788/0819 →
Continuity (1)
Related Publication 20230319292A1 · Oct 5, 2023
References Cited (47)
US 7225267B2 · Key et al. · 2007 [cited by applicant]
US 7916783B2 · Gao et al. · 2011 [cited by applicant]
US 9679258B2 · Mnih et al. · 2017 [cited by applicant]
US 10341670B1 · Brailovskiy et al. · 2019 [cited by applicant]
US 11330333B2 · Puente · 2022 [cited by examiner]
US 20060192850A1 · Verhaegh et al. · 2006 [cited by applicant]
US 20070025441A1 · Ugur et al. · 2007 [cited by applicant]
US 20140115100A1 · Changuel · 2014 [cited by examiner]
US 20140240319A1 · Yasser · 2014 [cited by applicant]
US 20160019303A1 · Littleford · 2016 [cited by applicant]
US 20210012227A1 · Fang · 2021 [cited by examiner]
CN 108063961 · 2018 [cited by applicant]
CN 108629422 · 2018 [cited by applicant]
CN 109769119 · 2019 [cited by applicant]
CN 111031387 · 2020 [cited by applicant]
JP 2001231039 · 2001 [cited by applicant]
JP 2006524461 · 2006 [cited by applicant]
RU 2579967 · 2016 [cited by applicant]
WO WO2019104635 · 2019 [cited by applicant]
WO WO2019117970 · 2019 [cited by applicant]
WO WO2020112321 · 2020 [cited by applicant]
Communication pursuant to Rules 161(2) and 162 EPC dated Feb. 7, 2023, from European Patent Application No. 20943454.7, 3 pp. [cited by applicant]
Communication pursuant to Rules 70(2) and 70a(2) EPC dated Mar. 11, 2024, from European Patent Application No. 20943454.7, 1 p. [cited by applicant]
Extended European Search Report dated Feb. 20, 2024, from European Patent Application No. 20943454.7, 10 pp. [cited by applicant]
Guo et al., “A Bayesian Approach to Block Structure Inference in AV1-based Multi-rate Video Encoding,” [cited by applicant]
Guo et al., “Rate Control for Screen Content Coding Based on Picture Classification,” [cited by applicant]
Guo et al., “Rate Control for Screen Content Coding in HEVC,” [cited by applicant]
Helle et al., “Reinforcement Learning for Video Encoder Control in HEVC,” [cited by applicant]
Hu et al., “Reinforcement Learning for HEVC/H.265 Intra-Frame Rate Control,” [cited by applicant]
Huang et al., “QARC: Video Quality Aware Rate Control for Real-Time Video Streaming via Deep Reinforcement Learning,” [cited by applicant]
Li et al., “A Convolutional Neural Network-Based Approach to Rate Control in HEVC Intra Coding,” [cited by applicant]
Official Action dated Nov. 8, 2023, from Russian Patent Application No. 2022133819, 18 pp. [cited by applicant]
Rippel et al., “Learned Video Compression,” [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms,” arXiv:1707.06347v2, 12 pp. (Aug. 2017). [cited by applicant]
Tang et al., “Down-Sampling Based Rate Control for Mobile Screen Video Coding,” [cited by applicant]
Zhang et al., “A New Rate Control Scheme for Video Coding Based on Region of Interest,” [cited by applicant]
Zheng et al., “Lightweight Content-Adaptive Coding in Joing Nalyzing-Encoding Framework,” [cited by applicant]
Zhou et al., “Rate Control Method Based on Deep Reinforcement Learning for Dynamic Video Sequences in Hevc,” [cited by applicant]
Zhu et al., “Low-Complexity Reinforcement Learning for Delay-Sensitive Compression in Networked Video Stream Mining,” [cited by applicant]
Office Action dated Feb. 2, 2024, from South African Patent Application No. 2022/11950, 1 p. [cited by applicant]
Official Action dated Mar. 28, 2024, from Russian Patent Application No. 2022133819, 10 pp. [cited by applicant]
Cheng et al., “Adaptive Video Transmission Control System Based on Reinforcement Learning Approach Over Heterogeneous Networks,” [cited by applicant]
International Search Report and Written Opinion dated Mar. 31, 2021, from International Patent Application No. PCT/CN2020/099390, 7 pp. [cited by applicant]
Kan, “Deep Reinforcement Learning-Based Rate Adaptation for Adaptive 360-Degree Video Streaming,” [cited by applicant]
Decision to Grant dated Aug. 1, 2024, from Russian Patent Application No. 2022133819, 17 pp. [cited by applicant]
Notice of Reasons for Refusal dated May 7, 2024, from Japanese Patent Application No. 2022-581327, 10 pp. [cited by applicant]
Notice of Reasons for Refusal dated Oct. 1, 2024, from Japanese Patent Application No. 2022-581327, 6 pp. [cited by applicant]