Affine motion prediction-based video decoding method and device using subblock-based temporal merge candidate in video coding system
A video decoding method performed by a decoding device according to the present document is characterized by including: a step for deriving reference subblocks in a reference picture on the basis of the motion vector of an adjacent block on the left side of the current block; a step for deriving a subblock-based temporal merge candidate for the current block on the basis of motion information about the reference subblocks; a step for forming an affine merge candidate list for the current block, the affine merge candidate list including the subblock-based temporal merge candidate; a step for deriving motion information about subblocks of the current block on the basis of the affine merge candidate list; a step for deriving prediction samples for the current block on the basis of the motion information about the subblocks; and a step for generating a reconstructed picture on the basis of the prediction samples.
1. A decoding apparatus for image decoding, the decoding apparatus comprising:
a memory; and
at least one processor connected to the memory, the at least one processor configured to:
derive a reference block in a reference picture based on a left neighboring block of a current block;
derive a subblock-based temporal merging candidate for the current block based on the reference block;
derive an inherited affine candidate for the current block;
derive a constructed affine candidate for the current block;
construct a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate;
derive motion information of sub-blocks of the current block based on the subblock merge candidate list;
derive prediction samples for the current block based on motion information of the sub-blocks; and
generate a reconstructed picture based on the prediction samples,
wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a y component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.
2. The decoding apparatus of claim 1 , wherein a position of the reference block is derived based on a motion vector of the left neighboring block.
3. The decoding apparatus of claim 2 , wherein the motion vector for deriving the position of the reference block is fixed to the motion vector of the left neighboring block.
4. The decoding apparatus of claim 1 , wherein the deriving the motion information of the sub-blocks of the current block comprises:
selecting the subblock-based temporal merging candidate from the subblock merge candidate list; and
deriving the motion information of the sub-blocks of the current block based on the subblock-based temporal merging candidate.
5. The decoding apparatus of claim 4 , wherein motion information of a target sub-block among the sub-blocks is derived based on motion information of a collocated sub-block for the target sub-block included in the subblock-based temporal merging candidate.
6. An encoding apparatus for image encoding, the encoding apparatus comprising:
a memory; and
at least one processor connected to the memory, the at least one processor configured to:
derive a reference block in a reference picture based on a left neighboring block of a current block;
derive a subblock-based temporal merging candidate for the current block based on motion information of the reference block;
derive an inherited affine candidate for the current block;
derive a constructed affine candidate for the current block;
construct a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate;
derive motion information of sub-blocks of the current block based on the subblock merge candidate list;
derive prediction information for the current block based on the motion information of the sub-blocks; and
encode image information including the prediction information for the current block,
wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a y component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.
7. The encoding apparatus of claim 6 , wherein a position of the reference block is derived based on a motion vector of the left neighboring block.
8. The encoding apparatus of claim 7 , wherein the motion vector for deriving the position of the reference block is fixed to the motion vector of the left neighboring block.
9. The encoding apparatus of claim 6 , wherein the deriving the motion information of the sub-blocks of the current block comprises:
selecting the subblock-based temporal merging candidate from the subblock merge candidate list; and
deriving the motion information of the sub-blocks of the current block based on the subblock-based temporal merging candidate.
10. The encoding apparatus of claim 9 , wherein motion information of a target sub-block among the sub-blocks is derived based on motion information of a collocated sub-block for the target sub-block included in the subblock-based temporal merging candidate.
11. A transmitting apparatus of data for an image, comprising:
at least one processor configured to obtain a bitstream for the image,
wherein the bitstream is generated based on deriving a reference block in a reference picture based on a left neighboring block of a current block, deriving a subblock-based temporal merging candidate for the current block based on motion information of the reference block, deriving an inherited affine candidate for the current block, deriving a constructed affine candidate for the current block, constructing a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate, deriving motion information of sub-blocks of the current block based on the subblock merge candidate list, deriving prediction information for the current block based on the motion information of the sub-blocks, encoding image information including the prediction information for the current block, and generating the bitstream including the image information; and
a transmitter configured to transmit the data comprising the bitstream,
wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a y component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.