IP Library › Granted Patent US 11,558,275
Granted Patent B2
US 11,558,275 · App. 16/877,257 · Granted Jan 17, 2023

Reinforcement learning for jitter buffer control

Inventors: Xiulian Peng (Beijing, CN); Vinod Prakash (Redmond, WA); Xiangyu Kong (Beijing, CN); Sriram Srinivasan (Sammamish, WA); Yan Lu (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
H04L43/087G06K9/6256G06K9/6262G06N20/00H04L41/145H04L47/283
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,558,275
App. No.
16/877,257
Granted
Jan 17, 2023
Kind
B2
Abstract

Disclosed in some examples are methods, systems, and machine-readable mediums which determine jitter buffer delay by inputting jitter buffer and currently observed network status information to a machine learned model that is trained using a reinforcement learning (RL) method. The model maps these inputs to an action to compress, stretch, or hold the jitter buffer delay, which is used by a recipient computing device to optimize the jitter buffer delay. The model may be trained using a simulator that uses network traces of past real streaming sessions (e.g., communication sessions) of users. By training the model through reinforcement learning, the model learns to make better decisions through reinforcement in the form of reward signals that reflect the performance of each decision.

Claims (53)

1. A device for controlling jitter-buffer delay in a media streaming session, the device comprising:

a computer processor;

a memory, storing instructions, which when executed by the computer processor causes the computer processor to perform operations comprising:

identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media data of an ongoing media streaming session taking place over a network, the media data stored in the jitter buffer prior to processing of the media data; the jitter buffer state comprising an indicator of a delay in processing of media data in the jitter buffer;

identifying a network delay of the network;

determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, the network delay, and a reward, the action is a stretch, compress, or hold action, and the reward is reduced based on each of a jitter buffer delay and whether a stretch action was performed as a most recent action; and

determining a playback duration of the media frame based upon the action.

2. The device of claim 1 , wherein the operations of determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises determining the action for the frame of media data based upon past jitter buffer states, past network states, past actions.

3. The device of claim 1 , wherein the jitter buffer state comprises a current jitter buffer delay, current received frames in the jitter buffer, total delay of the media frame, whether the media frame is concealed or not, whether the media frame is newly received, and a previously taken action.

4. The device of claim 1 , wherein the operations of determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises determining the action by using a model trained using past network traces of previous media streaming sessions.

5. The device of claim 4 , wherein the operations further comprise:

training the model by:

simulating a jitter buffer using the past network traces;

producing a training action using the model for the simulated jitter buffer;

producing an estimated value of the training action using a second model; and

modifying the model based upon the estimated value and the reward signal.

6. The device of claim 5 , wherein the model and second model share a layer.

7. The device of claim 4 , wherein the model comprises at least two layers wherein at least one layer includes a leaky rectified linear unit activation function.

8. The device of claim 4 , wherein the model comprises:

a first layer that comprises a leaky rectified linear unit (ReLu) activation function;

a second layer that comprises a leaky ReLu activation function;

a third layer that comprises a gated recurrent unit (GRU);

a fourth layer that comprises a leaky ReLu activation function; and

a fifth layer implementing a soft-max function.

9. The device of claim 1 , wherein the operations further comprise playing back the media frame at the playback duration.

10. The device of claim 1 , wherein the operations of determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises using a machine-learned model, and wherein the operations further comprise:

producing an estimated value of the action using a second model; and

modifying the model based upon the estimated value and the reward.

11. The device of claim 1 , wherein the media data comprises audio data, video data, or both audio and video data.

12. A method for controlling jitter-buffer delay in a media streaming session, the method comprising:

identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media data of an ongoing media streaming session taking place over a network, the media data stored in the jitter buffer prior to processing of the media data, the jitter buffer state comprising an indicator of a delay in processing of media data in the jitter buffer;

identifying a network delay of the network;

determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, the network delay, and a reward, the action is a stretch, compress, or hold action, and the reward is reduced based on each of a jitter buffer delay and whether a stretch action was performed as a most recent action; and

determining a playback duration of the media frame based upon the action.

13. The method of claim 12 , wherein determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises determining the action for the frame of media data based upon past jitter buffer states, past network states, past actions.

14. The method of claim 12 , wherein the jitter buffer state comprises a current jitter buffer delay, current received frames in the jitter buffer, total delay of the media frame, whether the media frame is concealed or not, whether the media frame is newly received, and a previously taken action.

15. The method of claim 12 , wherein determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises determining the action by using a model trained using past network traces of previous media streaming sessions.

16. The method of claim 15 , further comprising:

training the model by:

simulating a jitter buffer using the past network traces;

producing a training action using the model for the simulated jitter buffer;

producing an estimated value of the training action using a second model; and

modifying the model based upon the estimated value and the reward-signal.

17. The method of claim 12 , wherein determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises using a machine-learned model, and wherein the method further comprises:

producing an estimated value of the action using a second model; and

modifying the model based upon the estimated value and the reward.

18. The method of claim 12 , wherein the media data comprises audio data, video data, or both audio and video data.

19. A device for controlling jitter-buffer delay in a media streaming session, the device comprising:

means for identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media data of an ongoing media streaming session taking place over a network, the media data stored in the jitter buffer prior to processing of the media data, the jitter buffer state comprising an indicator of a delay in processing of media data in the jitter buffer;

means for identifying a network delay of the network;

means for determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, and the network delay, and a reward, the action is a stretch, compress, or hold action, and the reward is reduced based on each of a jitter buffer delay and whether a stretch action was performed as a most recent action; and

means for determining a playback duration of the media frame based upon the action.

20. The device of claim 19 , wherein the means for determining the action for the frame of media data in the jitter buffer based upon the jitter buffer state and network delay comprises means for determining the action for the frame of media data based upon past jitter buffer sates, past network states, past actions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2020
From: PENG, XIULIAN; PRAKASH, VINOD; KONG, XIANGYU; SRINIVASAN, SRIRAM; LU, YAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 052758/0198 →
Continuity (2)
Provisional Application 62976047 · Feb 13, 2020
Related Publication 20210258235A1 · Aug 19, 2021