Transform coding based on matrix-based intra prediction
Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to a rule, where the rule specifies a relationship between applicability of a matrix based intra prediction (MIP) mode or a transform mode during the conversion, where the MIP mode includes determining a prediction block of the current video block by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation, and where the transform mode specifies use of a transform operation for the determining the prediction block for the current video block.
1. A video processing method, comprising:
determining, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,
performing a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and
performing the conversion based on the prediction samples and residual samples of the video block;
wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and
wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.
2. The method of claim 1 , wherein the secondary transform tool includes a non-separable secondary transform.
3. The method of claim 1 , wherein the secondary transform tool includes a Reduced Secondary Transform (RST) or a rotation transform.
4. The method of claim 1 , wherein based on the first intra mode being applied on the video block, a second intra mode is converted from the first intra mode, and wherein a secondary transform classification is derived based on the second intra mode.
5. The method of claim 4 , wherein the selected secondary transform matrix is determined based on the secondary transform classification.
6. The method of claim 1 , wherein a down-sampling factor derived in the boundary down-sampling operation is larger than or equal to 1, and wherein a one-dimensional vector array is further derived based on concatenating down-sampled reference samples derived from the boundary down-sampling operation and the one-dimensional vector array is used as input of the matrix vector multiplication operation.
7. The method of claim 4 , wherein in the second intra mode, distance-based weighted calculations are applied on reference values in vertical direction and horizontal direction to derive prediction values.
8. The method of claim 4 , wherein the second intra mode includes a planar mode.
9. The method of claim 1 , wherein whether to apply the first intra mode is specified by a first syntax element included in a sequence parameter set and a second syntax element included in a coding unit level set.
10. The method of claim 9 , wherein in response to a width-height ratio of the video block being greater than 2, a context with an index of 3 is used for a first bin of the second syntax element.
11. The method of claim 9 , wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a single context selected from contexts with indices of 0, 1 or 2 is used for a first bin of the second syntax element.
12. The method of claim 1 , wherein the first intra mode includes multiple types, and a type index for the video block is derived excluding referring to type indices of previous video blocks.
13. The method of claim 12 , wherein the type index for the video block is explicitly included in the bitstream.
14. The method of claim 1 , wherein whether to apply the secondary transform tool is based on a height (H) or a width (W) of the video block.
15. The method of claim 1 , wherein the conversion includes encoding the video block into the bitstream.
16. The method of claim 1 , wherein the conversion includes decoding the video block from the bitstream.
17. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:
determine, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,
perform a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and
perform the conversion based on the prediction samples and residual samples of the video block;
wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and
wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.
18. The apparatus of claim 17 , wherein based on the first intra mode being applied on the video block, a second intra mode is converted from the first intra mode, and wherein a secondary transform classification is derived based on the second intra mode.
19. A non-transitory computer-readable storage medium storing instructions that cause a processor to:
determine, for a conversion between a video block of a video and the bitstream of the video, that a first intra mode is applied on the video block of the video,
perform a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and
perform the conversion based on the prediction samples and residual samples of the video block;
wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and
wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.
20. A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:
determining that a first intra mode is applied on a video block of the video,
performing a boundary downsampling operation on reference samples of the video block based on a size of the video block, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation to generate prediction samples for the video block; and
generating the bitstream based on the prediction samples and residual samples of the video block;
wherein a secondary transform tool is applied on the video block, and a secondary transform matrix is selected to generate the residual samples for the video block, and
wherein the secondary transform tool includes applying, during encoding, a forward secondary transform to an output of a forward primary transform before applying a quantization process, or applying, during decoding, an inverse secondary transform to an output of a dequantization process before applying an inverse primary transform.