Neural network based coefficient sign prediction field
A method, computer program, and computer system is provided for coding video data. Reference samples and magnitudes of transform coefficients corresponding to a current block of video data from an input to a neural network are identified. Sign values associated with the transform coefficients are predicted using neural networks. The video data is encoded/decoded based on the predicted sign values.
1. A method of video decoding, executable by a processor, the method comprising:
receiving a video bitstream comprising a current block in a current picture of video data of the video bitstream;
identifying reference samples and magnitudes of transform coefficients, corresponding to the current block, from an input to a neural network;
predicting sign values associated with the transform coefficients based at least on the identified reference samples, wherein the sign values are predicted using a convolutional neural network; and
decoding the video data based on the predicted sign values.
2. The method of claim 1 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.
3. The method of claim 1 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.
4. The method of claim 1 , wherein the transform coefficients are de-quantized transform coefficients.
5. The method of claim 1 , wherein the predicted sign values for correspond to a limited set of transform coefficients.
6. The method of claim 5 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.
7. The method of claim 5 , wherein the limited set of transform coefficients includes a pre-defined number of coefficients along a forward scanning order.
8. A method of video encoding, executable by a processor, the method comprising:
receiving video data;
identifying reference samples and magnitudes of transform coefficients, corresponding to a current block of a picture of the video data, from an input to a neural network;
determining sign values associated with the transform coefficients based at least on the identified reference samples, wherein the sign values are predicted using a convolutional neural network; and
encoding the video data based on the determined sign values.
9. The method of claim 8 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.
10. The method of claim 8 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.
11. The method of claim 8 , wherein the transform coefficients are de-quantized transform coefficients.
12. The method of claim 8 , wherein the predicted sign values for correspond to a limited set of transform coefficients.
13. The method of claim 12 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.
14. The method of claim 12 , wherein the limited set of transform coefficients includes a pre-defined number of coefficients along a forward scanning order.
15. A method of processing visual media data, executable by a processor, the method comprising:
performing a conversion between a visual media and a bitstream of visual media data according to a format rule, wherein the bitstream comprises video data including a current block in a current picture, and wherein the format rule specifies:
identifying reference samples and magnitudes of transform coefficients, corresponding to the current block of the current picture, from an input to a neural network, and
predicting sign values associated with the transform coefficients using a convolutional neural network based at least on the identified reference samples.
16. The method of claim 15 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.
17. The method of claim 15 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.
18. The method of claim 15 , wherein the transform coefficients are de-quantized transform coefficients.
19. The method of claim 15 , wherein the predicted sign values for correspond to a limited set of transform coefficients.
20. The method of claim 19 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.