End-to-end dependent quantization with deep reinforcement learning
There is included a method and apparatus comprising computer code configured to cause a processor or processors to perform obtaining an input stream of video data, computing a key based on a floating number in the input stream, predicting a current dependent quantization (DQ) state based on a state predictor and a number of previous keys and a number of previous DQ states, reconstructing the floating number based on the key and the current DQ state, and coding the video based on the reconstructed floating number.
1. A method for video coding performed by at least one processor, the method comprising:
obtaining an input stream of video data;
computing a key based on a floating number in the input stream;
predicting a current dependent quantization (DQ) state based on a state predictor and a number of previous keys and a number of previous DQ states;
reconstructing the floating number based on the key and the current DQ state; and
coding the video based on the reconstructed floating number.
2. The method according to claim 1 ,
wherein computing the key and reconstructing the floating number comprises implementing one or more deep neural networks (DNN).
3. The method according to claim 1 ,
wherein the state predictor comprises an action-value mapping function between an action and an output Q-value associated with the action.
4. The method according to claim 3 , further comprising:
computing a plurality of keys, including the key, based on a plurality of floating numbers, including the floating number, in the input stream; and
reconstructing the plurality of floating numbers based on the plurality of keys and at least the current DQ state.
5. The method according to claim 3 ,
wherein the action corresponds to at least one of the DQ states.
6. The method according to claim 5 , further comprising:
wherein the state predictor further comprises respective correspondences between ones of a plurality of actions, including the action, and ones of the DQ states, including the at least one of the DQ states.
7. The method according to claim 1 ,
wherein predicting the current DQ state comprises implementing an action-value mapping function between an action and an output Q-value associated with the action the previous keys and the previous DQ states.
8. The method according to claim 1 ,
wherein the state predictor comprises an action-value mapping function between an action and an output Q-value associated with the action, and
wherein the output Q-value represents a measurement of a target quantization performance associated with a sequence of actions, including the action.
9. The method according to claim 1 ,
wherein predicting the current DQ state based on the state predictor comprises computing Q-values, including the output Q-value, for each of the actions.
10. The method according to claim 1 ,
wherein the output Q-value is selected from among the computed Q-values.
11. An apparatus for video coding performed by at least one processor, the apparatus comprising:
at least one memory configured to store computer program code;
at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including:
obtaining code configured to cause the at least one processor to obtain an input stream of video data;
computing code configured to cause the at least one processor to compute a key based on a floating number in the input stream;
predicting code configured to cause the at least one processor to predict a current dependent quantization (DQ) state based on a state predictor and a number of previous keys and a number of previous DQ states;
reconstructing code configured to cause the at least one processor to reconstruct the floating number based on the key and the current DQ state; and
coding code configured to cause the at least one processor to code the video based on the reconstructed floating number.
12. The apparatus according to claim 11 ,
wherein computing the key and reconstructing the floating number comprises implementing one or more deep neural networks (DNN).
13. The apparatus according to claim 11 ,
wherein the state predictor comprises an action-value mapping function between an action and an output Q-value associated with the action.
14. The apparatus according to claim 13 ,
wherein the computing code is further configured to cause the at least one processor to compute a plurality of keys, including the key, based on a plurality of floating numbers, including the floating number, in the input stream; and
wherein the reconstructing code is further configured to cause the at least one processor to reconstruct the plurality of floating numbers based on the plurality of keys and at least the current DQ state.
15. The apparatus according to claim 14 ,
wherein the action corresponds to at least one of the DQ states.
16. The apparatus according to claim 15 , wherein the state predictor further comprises respective correspondences between ones of a plurality of actions, including the actions, and ones of the DQ states, including the at least one of the DQ states.
17. The apparatus according to claim 1 ,
wherein predicting the current DQ state comprises implementing an action-value mapping function between an action and an output Q-value associated with the action the previous keys and the previous DQ states.
18. The apparatus according to claim 1 ,
wherein the state predictor comprises an action-value mapping function between an action and an output Q-value associated with the action, and
wherein the output Q-value represents a measurement of a target quantization performance associated with a sequence of actions, including the action.
19. The apparatus according to claim 1 ,
wherein predicting the current DQ state based on the state predictor comprises computing Q-values, including the output Q-value, for each of the actions, and
wherein the output Q-value is selected from among the computed Q-values.
20. A non-transitory computer readable medium storing a program causing a computer to execute a process, the process comprising:
obtaining an input stream of video data;
computing a key based on a floating number in the input stream;
predicting a current dependent quantization (DQ) state based on a state predictor and a number of previous keys and a number of previous DQ states;
reconstructing the floating number based on the key and the current DQ state; and
coding the video based on the reconstructed floating number.