Matrix-based intra prediction using filtering
Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix based intra prediction (MIP) mode in which a prediction block of the current video block is determined by performing, on reference boundary samples located to a left of the current video block and located to a top of the current video block, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation, where instead of reduced boundary samples calculated from the reference boundary samples of the current video block in the boundary downsampling operation, the reference boundary samples are directly used for a prediction process in the upsampling operation.
1 . A method of processing video data, comprising:
determining, for a conversion between a video block of a video and a bitstream of the video, that a first intra mode is applied on the video block of the video, wherein process in the first intra mode includes a one-stage boundary downsampling operation, followed by a matrix vector multiplication operation and selectively followed by an upsampling operation to generate prediction samples for the video block of the video; and
performing the conversion based on the prediction samples;
wherein, in the one-stage boundary downsampling operation, when a downsampled boundary size of the video block is less than a size of the of the video block, reduced boundary samples are generated from reference boundary samples of the video block of the video by the one-stage boundary downsampling operation and are used to generate inputs to the matrix vector multiplication operation, and the reduced boundary samples are generated directly from the reference boundary samples of the video block and a downscaling factor without deriving intermediate samples,
wherein the downscaling factor is calculated only once for a horizontal direction and a vertical direction, respectively,
wherein the reference boundary samples include left and above reference boundary samples of the video block,
wherein a first syntax element indicating whether to apply the first intra mode is included in the bitstream, wherein at least one bin of the first syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the video block, wherein the increasement value of the context is determined further based on a size of the video block,
wherein in response to a width-height ratio of the video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the first syntax element, and
wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the first syntax element, wherein the second increasement value is not identical to the first predefined increasement value.
2 . The method of claim 1 , wherein the left and above reference boundary samples are derived without an intra reference sample filtering process.
3 . The method of claim 1 , wherein the downscaling factor is calculated based on a size of the video block.
4 . The method of claim 1 , wherein inputting samples to the upsamping operation include upsampling boundary samples of the video block, and the upsampling boundary samples are not computed by averaging the reference boundary samples of the video block.
5 . The method of claim 4 , wherein at least one of the upsampling boundary samples are copied from the reference boundary samples.
6 . The method of claim 1 , wherein the matrix vector multiplication operation is followed by a transposing operation prior to the upsampling operation.
7 . The method of claim 6 , wherein the transposing operation converts a block having a width of a first value and a height of a second value to a block having the width of the second value and a height of the first value.
8 . The method of claim 1 , wherein the conversion includes encoding the video block into the bitstream.
9 . The method of claim 1 , wherein the conversion includes decoding the video block from the bitstream.
10 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:
determine, for a conversion between a video block of a video and a bitstream of the video, that a first intra mode is applied on the video block of the video, wherein process in the first intra mode includes a one-stage boundary downsampling operation, followed by a matrix vector multiplication operation and selectively followed by an upsampling operation to generate prediction samples for the video block of the video; and
perform the conversion based on the prediction samples;
wherein, in the one-stage boundary downsampling operation, when a downsampled boundary size of the video block is less than a size of the of the video block, reduced boundary samples are generated from reference boundary samples of the video block of the video by the one-stage boundary downsampling operation and are used to generate inputs to the matrix vector multiplication operation, and the reduced boundary samples are generated directly from the reference boundary samples of the video block and a downscaling factor without deriving intermediate samples,
wherein the downscaling factor is calculated only once for a horizontal direction and a vertical direction, respectively,
wherein the reference boundary samples include left and above reference boundary samples of the video block,
wherein a first syntax element indicating whether to apply the first intra mode is included in the bitstream, wherein at least one bin of the first syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the video block, wherein the increasement value of the context is determined further based on a size of the video block,
wherein in response to a width-height ratio of the video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the first syntax element, and
wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the first syntax element, wherein the second increasement value is not identical to the first predefined increasement value.
11 . The apparatus of claim 10 , wherein the left and above reference boundary samples of the video block are derived without an intra reference sample filtering process.
12 . The apparatus of claim 10 , wherein the downscaling factor is calculated based on a size of the video block.
13 . The apparatus of claim 10 , wherein inputting samples to the upsamping operation include upsampling boundary samples of the video block, wherein at least one of the upsampling boundary samples are copied from the reference boundary samples.
14 . The apparatus of claim 10 , wherein inputting samples to the upsamping operation include upsampling boundary samples of the video block, and the upsampling boundary samples are not computed by averaging the reference boundary samples of the video block.
15 . The apparatus of claim 10 , wherein the matrix vector multiplication operation is followed by a transposing operation prior to the upsampling operation, and wherein the transposing operation converts a block having a width of a first value and a height of a second value to a block having the width of the second value and a height of the first value.
16 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:
determine, for a conversion between a video block of a video and a bitstream of the video, that a first intra mode is applied on the video block of the video, wherein process in the first intra mode includes a one-stage boundary downsampling operation, followed by a matrix vector multiplication operation and selectively followed by an upsampling operation to generate prediction samples for the video block of the video; and
perform the conversion based on the prediction samples;
wherein, in the one-stage boundary downsampling operation, when a downsampled boundary size of the video block is less than a size of the of the video block, reduced samples are generated from reference boundary samples of the video block of the video by the one-stage boundary downsampling operation and are used to generate inputs to the matrix vector multiplication operation, and the reduced boundary samples are generated directly from the reference boundary samples of the video block and a downscaling factor without deriving intermediate samples,
wherein the downscaling factor is calculated only once for a horizontal direction and a vertical direction, respectively,
wherein the reference boundary samples include left and above reference boundary samples of the video block,
wherein a first syntax element indicating whether to apply the first intra mode is included in the bitstream, wherein at least one bin of the first syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the video block, wherein the increasement value of the context is determined further based on a size of the video block,
wherein in response to a width-height ratio of the video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the first syntax element, and
wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the first syntax element, wherein the second increasement value is not identical to the first predefined increasement value.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the left and above reference boundary samples are derived without an intra reference sample filtering process.
18 . The non-transitory computer-readable storage medium of claim 16 , wherein the matrix vector multiplication operation is followed by a transposing operation prior to the upsampling operation, and wherein the transposing operation converts a block having a width of a first value and a height of a second value to a block having the width of the second value and a height of the first value.
19 . A method for storing a bitstream of a video, comprising:
determining, that a first intra mode is applied on a video block of the video, wherein process in the first intra mode includes a one-stage boundary downsampling operation, followed by a matrix vector multiplication operation and selectively followed by an upsampling operation to generate prediction samples for the video block of the video; and
generating the bitstream based on the determining;
wherein, in the one-stage boundary downsampling operation, when a downsampled boundary size of the video block is less than a size of the of the video block, reduced samples are generated from reference boundary samples of the video block of the video by the one-stage boundary downsampling operation and are used to generate inputs to the matrix vector multiplication operation, and the reduced boundary samples are generated directly from the reference boundary samples of the video block and a downscaling factor without deriving intermediate samples,
wherein the downscaling factor is calculated only once for a horizontal direction and a vertical direction, respectively,
wherein the reference boundary samples include left and above reference boundary samples of the video block,
wherein a first syntax element indicating whether to apply the first intra mode is included in the bitstream, wherein at least one bin of the first syntax element is context coded, and an increasement value of the context is determined based on characteristics of a neighboring block of the video block, wherein the increasement value of the context is determined further based on a size of the video block,
wherein in response to a width-height ratio of the video block being greater than 2, a context with a first predefined increasement value is used for coding the at least one bin of the first syntax element, and
wherein in response to a width-height ratio of the video block being smaller than or equal to 2, a context with a second increasement value is used for coding the at least one bin of the first syntax element, wherein the second increasement value is not identical to the first predefined increasement value.
20 . The method of claim 19 , wherein the matrix vector multiplication operation is followed by a transposing operation prior to the upsampling operation, and wherein the transposing operation converts a block having a width of a first value and a height of a second value to a block having the width of the second value and a height of the first value.