Neural network based transform set selection
Systems and methods for coding and decoding of a coded bitstream is provided. A method includes encoding a block of a picture. The encoding includes selecting a transform set based on at least one neighboring sample from one or more previously encoded neighboring blocks or from a previously encoded picture and transforming coefficients of the block using a transform from the transform set.
1 . A method performed by at least one processor, the method comprising:
decoding a block of a picture, the decoding comprising:
selecting a transform set based on at least one neighboring sample from one or more neighboring blocks or from a picture, and
the at least one neighboring sample is an input to a neural network whose output is a selection of the transform set and an identifier of a prediction mode set; and
transforming coefficients of the block using a transform from the transform set, and the selecting the transform set comprises:
selecting a sub-group of transform sets from a group of transform sets based on information of an intra prediction mode or an inter prediction mode; and
selecting the transform set from the sub-group.
2 . The method of claim 1 , wherein the neural network identifies the transform set as for at least one of a secondary transform, a primary transform, and a combination of the secondary transform and the secondary transform.
3 . The method of claim 2 , wherein the secondary transform indicates a non-separable transform scheme, and wherein the primary transform indicates any of different types of discrete cosine transforms, discrete sine transforms, line graph transforms with respectively different self-loop rates.
4 . The method of claim 1 , wherein the at least one neighboring sample is both at least one of upsampled and downsampled and also an input to the neural network.
5 . The method of claim 1 , wherein the at least one neighboring sample is scaled and input to the neural network.
6 . The method of claim 1 , wherein parameters of the neural network are based on any of whether the block is intra coded, a block width, a block height, a quantization parameter, whether a current picture is coded as an intra frame, and an intra prediction mode.
7 . The method of claim 1 , wherein the at least one neighboring sample includes one or more lines of top and left neighboring reconstructed samples.
8 . A method performed by at least one processor, the method comprising: encoding a block of a picture, the encoding comprising:
selecting a transform set based on at least one neighboring sample from one or more neighboring blocks or from a picture, and the at least one neighboring sample is an input to a neural network whose output is a selection of the transform set and an identifier of a prediction mode set; and
transforming coefficients of the block using a transform from the transform set, and the selecting the transform set comprises:
selecting a sub-group of transform sets from a group of transform sets based on information of an intra prediction mode or an inter prediction mode; and
selecting the transform set from the sub-group.
9 . The method of claim 8 , wherein the neural network identifies the transform set as for at least one of a secondary transform, a primary transform, and a combination of the secondary transform and the secondary transform.
10 . The method of claim 9 , wherein the secondary transform indicates a non-separable transform scheme, and wherein the primary transform indicates any of different types of discrete cosine transforms, discrete sine transforms, line graph transforms with respectively different self-loop rates.
11 . The method of claim 8 , wherein the at least one neighboring sample is both at least one of upsampled and downsampled and also an input to the neural network.
12 . The method of claim 8 , wherein the at least one neighboring sample is scaled and input to the neural network.
13 . The method of claim 8 , wherein parameters of the neural network are based on any of whether the block is intra coded, a block width, a block height, a quantization parameter, whether a current picture is coded as an intra frame, and an intra prediction mode.
14 . The method of claim 8 , wherein the at least one neighboring sample includes one or more lines of top and left neighboring reconstructed samples.
15 . A non-transitory computer-readable storage medium storing a video bitstream that is generated by a video encoding method, the video encoding method comprising: encoding and transmitting a block of a picture, the encoding comprising:
selecting a transform set based on at least one neighboring sample from one or more neighboring blocks or from a picture, and the at least one neighboring sample is an input to a neural network whose output is a selection of the transform set and an identifier of a prediction mode set; and
transforming coefficients of the block using a transform from the transform set, and the selecting the transform set comprises:
selecting a sub-group of transform sets from a group of transform sets based on information of an intra prediction mode or an inter prediction mode; and
selecting the transform set from the sub-group.
16 . The method of claim 15 , wherein the neural network identifies the transform set as for at least one of a secondary transform, a primary transform, and a combination of the secondary transform and the secondary transform.
17 . The method of claim 16 , wherein the secondary transform indicates a non-separable transform scheme, and wherein the primary transform indicates any of different types of discrete cosine transforms, discrete sine transforms, line graph transforms with respectively different self-loop rates.
18 . The method of claim 15 , wherein the at least one neighboring sample is both at least one of upsampled and downsampled and also an input to the neural network.
19 . The method of claim 15 , wherein the at least one neighboring sample is scaled and input to the neural network.