Signaling of coding tree unit block partitioning in neural network model compression
A method of neural network decoding includes receiving a first syntax element in a model parameter set from a bitstream of a compressed neural network representation (NNR) of a neural network. The first syntax element indicates whether a coding tree unit (CTU) block partitioning is enabled for a tensor in an NNR aggregate unit. The method also includes reconstructing the tensor in the NNR aggregate unit based on the first syntax element.
1. A method of neural network decoding at a decoder, comprising:
receiving a first syntax element in a neural network representation (NNR) aggregate unit header of an NNR aggregate unit from a bitstream of a compressed NNR of a neural network, the first syntax element indicating a maximum bit depth of quantized coefficients for a tensor in the NNR aggregate unit; and
reconstructing the tensor in the NNR aggregate unit based on the first syntax element.
2. The method of claim 1 , comprising:
receiving a third syntax element in a model parameter set from the bitstream of the compressed NNR of the neural network, the third syntax element indicating whether a coding tree unit (CTU) block partitioning is enabled for the tensor in the NNR aggregate unit; and
reconstructing the tensor in the NNR aggregate unit based on the third syntax element.
3. The method of claim 2 , wherein the third syntax element is (i) a model-wise syntax element to specify whether the CTU block partitioning is enabled for layers of the neural network, or (ii) a tensor-wise syntax element to specify whether the CTU block partitioning is enabled for the tensor in the NNR aggregate unit.
4. The method of claim 1 , further comprising:
receiving a second syntax element in the NNR aggregate unit header of the NNR aggregate unit from the bitstream, the second syntax element indicating a coding tree unit (CTU) scan order for processing the tensor in the NNR aggregate unit.
5. The method of claim 4 , wherein a first value of the second syntax element indicates that the CTU scan order is a first raster scan order at a horizontal direction, and a second value of the second syntax element indicates that the CTU scan order is a second raster scan order at a vertical direction.
6. The method of claim 1 , further comprising:
receiving a model-wise or tensor-wise fourth syntax element indicating a CTU dimension for the tensor in the NNR aggregate unit.
7. The method of claim 1 , further comprising:
receiving an NNR unit before receiving any NNR aggregate units, the NNR unit including a fifth syntax element indicating whether CTU partitioning is enabled.
8. A neural network encoding apparatus, comprising:
processing circuitry configured to
encode a first syntax element in a neural network representation (NNR) aggregate unit header of an NNR aggregate unit of a compressed NNR of a neural network, the first syntax element indicating a maximum bit depth of quantized coefficients for a tensor in the NNR aggregate unit unit; and
generate a bitstream of the compressed NNR of the neural network including the encoded first syntax element.
9. The apparatus of claim 8 , wherein the processing circuitry is further configured to:
encode a third syntax element in a model parameter set of the bitstream of the compressed NNR of the neural network, the third syntax element indicating whether a coding tree unit (CTU) block partitioning is enabled for the tensor in the NNR aggregate unit.
10. The apparatus of claim 9 , wherein the third syntax element is (i) a model-wise syntax element to specify whether the CTU block partitioning is enabled for layers of the neural network, or (ii) a tensor-wise syntax element to specify whether the CTU block partitioning is enabled for a tensor in the NNR aggregate unit.
11. The apparatus of claim 8 , wherein the processing circuitry is further configured to
encode a second syntax element in the NNR aggregate unit header of the NNR aggregate unit in the bitstream, the second syntax element indicating a coding tree unit (CTU) scan order for processing a tensor in the NNR aggregate unit.
12. The apparatus of claim 11 , wherein a first value of the second syntax element indicates that the CTU scan order is a first raster scan order at a horizontal direction, and a second value of the second syntax element indicates that the CTU scan order is a second raster scan order at a vertical direction.
13. The apparatus of claim 8 , wherein the processing circuitry is further configured to
encode a model-wise or tensor-wise fourth syntax element indicating a CTU dimension for a tensor in the NNR aggregate unit.
14. The apparatus of claim 8 , wherein the processing circuitry is further configured to
encode an NNR unit before encoding any NNR aggregate units, the NNR unit including a fifth syntax element indicating whether CTU partitioning is enabled.
15. A method of processing visual media data, the method comprising:
processing a bitstream of the visual media data according to a format rule, wherein
the bitstream includes a first syntax element in a neural network representation (NNR) aggregate unit header of an NNR aggregate unit of a compressed NNR of a neural network, the first syntax element indicating a maximum bit depth of quantized coefficients for a tensor in the NNR aggregate unit,
the format rule specifies that the tensor in the NNR aggregate unit is reconstructed based on the first syntax element.
16. The method of claim 15 , comprising:
receiving a third syntax element in a model parameter set from the bitstream of the compressed NNR of the neural network, the third syntax element indicating whether a coding tree unit (CTU) block partitioning is enabled for the tensor in the NNR aggregate unit; and
reconstructing the tensor in the NNR aggregate unit based on the third syntax element.
17. The method of claim 16 , wherein the third syntax element is (i) a model-wise syntax element to specify whether the CTU block partitioning is enabled for layers of the neural network, or (ii) a tensor-wise syntax element to specify whether the CTU block partitioning is enabled for the tensor in the NNR aggregate unit.
18. The method of claim 15 , further comprising:
receiving a second syntax element in the NNR aggregate unit header of the NNR aggregate unit from the bitstream, the second syntax element indicating a coding tree unit (CTU) scan order for processing the tensor in the NNR aggregate unit.
19. The method of claim 18 , wherein a first value of the second syntax element indicates that the CTU scan order is a first raster scan order at a horizontal direction, and a second value of the second syntax element indicates that the CTU scan order is a second raster scan order at a vertical direction.
20. The method of claim 15 , further comprising:
receiving a model-wise or tensor-wise fourth syntax element indicating a CTU dimension for the tensor in the NNR aggregate unit.