IP Library › Granted Patent US 12,113,680
Granted Patent B2
US 12,113,680 · App. 18/091,992 · Granted Oct 8, 2024

Reinforcement learning for jitter buffer control

Inventors: Xiulian Peng (Beijing, CN); Vinod Prakash (Redmond, WA); Xiangyu Kong (Beijing, CN); Sriram Srinivasan (Sammamish, WA); Yan Lu (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
H04L41/16G06F18/214G06F18/217G06N20/00H04L41/145H04L43/087H04L47/283
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,113,680
App. No.
18/091,992
Granted
Oct 8, 2024
Kind
B2
Abstract

Disclosed in some examples are methods, systems, and machine-readable mediums which determine jitter buffer delay by inputting jitter buffer and currently observed network status information to a machine learned model that is trained using a reinforcement learning (RL) method. The model maps these inputs to an action to compress, stretch, or hold the jitter buffer delay, which is used by a recipient computing device to optimize the jitter buffer delay. The model may be trained using a simulator that uses network traces of past real streaming sessions (e.g., communication sessions) of users. By training the model through reinforcement learning, the model learns to make better decisions through reinforcement in the form of reward signals that reflect the performance of each decision.

Claims (53)

1. A device for controlling jitter-buffer delay in a media streaming session over a network, the device comprising:

a computer processor;

a memory, storing instructions, which when executed by the computer processor causes the computer processor to perform operations comprising:

identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media frames of a media streaming session, the jitter buffer state comprising a current jitter buffer delay, current received frames in the jitter buffer, total delay of a current media frame of the media frames, and an immediately previous action;

identifying a network delay of the network;

determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, the network delay, and a reward; and

determining a playback duration of a next media frame based upon the action.

2. The device of claim 1 , wherein the operations of determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises determining the action for the media frame based upon past jitter buffer states, past network states, and past actions.

3. The device of claim 1 , wherein the jitter buffer state further comprises, whether the media frame is concealed or not, and whether the media frame is newly received.

4. The device of claim 1 , wherein the operations of determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises determining the action by using a model trained using past network traces of previous media streaming sessions.

5. The device of claim 4 , wherein the operations further comprise:

training the model by:

simulating a jitter buffer using the past network traces;

producing a training action using the model for the simulated jitter buffer;

producing an estimated value of the training action using a second model; and

modifying the model based upon the estimated value and the reward.

6. The device of claim 5 , wherein the model and second model share a neural network layer.

7. The device of claim 4 , wherein the model comprises at least two layers wherein at least one layer includes a leaky rectified linear unit activation function.

8. The device of claim 4 , wherein the model comprises:

a first layer that comprises a leaky rectified linear unit (ReLu) activation function;

a second layer that comprises a leaky ReLu activation function;

a third layer that comprises a gated recurrent unit (GRU);

a fourth layer that comprises a leaky ReLu activation function; and

a fifth layer implementing a soft-max function.

9. The device of claim 1 , wherein the operations further comprise playing back the media frame at the playback duration.

10. The device of claim 1 , wherein the operations of determining the action for the media frame based upon the jitter buffer state and the network delay comprises using a machine-learning model, and wherein the operations further comprise:

producing an estimated value of the action using a second model; and

modifying the model based upon the estimated value and the reward.

11. The device of claim 1 , wherein the media frame comprises audio data, video data, or both audio and video data.

12. A method for controlling jitter-buffer delay in a media streaming session over a network, the method comprising:

identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media frames of a media streaming session, the jitter buffer state comprising a current jitter buffer delay, current received frames in the jitter buffer, total delay of a current media frame of the media frames, and an immediately previous action;

identifying a network delay of the network;

determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, the network delay, and a reward; and

determining a playback duration of a next media frame based upon the action.

13. The method of claim 12 , wherein determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises determining the action for the media frame based upon past jitter buffer states, past network states, and past actions.

14. The method of claim 12 , wherein the jitter buffer state further comprises whether the media frame is concealed or not and whether the media frame is newly received.

15. The method of claim 12 , wherein determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises determining the action by using a model trained using past network traces of previous media streaming sessions.

16. The method of claim 15 , further comprising:

training the model by:

simulating a jitter buffer using the past network traces;

producing a training action using the model for the simulated jitter buffer;

producing an estimated value of the training action using a second model; and

modifying the model based upon the estimated value and the reward.

17. The method of claim 12 , wherein determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises using a machine-learned model, and wherein the method further comprises:

producing an estimated value of the action using a second model; and

modifying the model based upon the estimated value and the reward.

18. The method of claim 12 , wherein the media frame comprises audio data, video data, or both audio and video data.

19. A device for controlling jitter-buffer delay in a media streaming session over a network, the device comprising:

means for identifying a jitter buffer state of a jitter buffer, the jitter buffer storing media frames of a media streaming session, the jitter buffer state comprising a current jitter buffer delay, current received frames in the jitter buffer, total delay of a current media frame of the media frames, and an immediately previous action;

means for identifying a network delay of the network;

means for determining an action for a media frame of media data in the jitter buffer based upon the jitter buffer state, the network delay, and a reward; and

means for determining a playback duration of a next media frame based upon the action.

20. The device of claim 19 , wherein the means for determining the action for the media frame in the jitter buffer based upon the jitter buffer state and the network delay comprises means for determining the action for the media frame based upon past jitter buffer states, past network states, and past actions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: PENG, XIULIAN; PRAKASH, VINOD; KONG, XIANGYU; SRINIVASAN, SRIRAM; LU, YAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063001/0932 →
Continuity (3)
Continuation 16877257 · May 18, 2020
Provisional Application 62976047 · Feb 13, 2020
Related Publication 20230138038A1 · May 4, 2023