Multi-iteration motion vector refinement
A method for video processing includes: refining motion information of a video block by using a multi-step refinement processing, multiple refined motion vectors (MVs) of the video block being derived iteratively in respective steps of the multi-step refinement processing, and performing a video processing on the video block based on the multiple refined MVs of the video block.
1. A method for video processing, comprising:
refining motion information of a video block by using a multi-step refinement processing, wherein multiple refined motion vectors (MVs) of the video block are derived iteratively in respective steps of the multi-step refinement processing, and
performing a video processing on the video block based on the multiple refined MVs of the video block;
wherein the refining comprises:
acquiring original MVs (MVLX0_x, MVLX0_y), to generate at least one original motion compensated reference block, wherein LX=L0 or L1, where L0 and L1 represent reference list 0 and list 1 respectively,
refining, based on at least one original motion compensated reference block, the original MVs (MVLX0_x, MVLX0_y) to derive first refined MVs (MVLX1_x, MVLX1_y), to generate at least one first motion compensated reference block, and
refining the first refined MVs (MVLX1_x, MVLX1_y) to derive second refined MVs (MVLX2_x, MVLX2_y), wherein the first motion compensated reference block is used to derive at least one of temporal gradients and spatial gradients used for deriving the second refined MVs (MVLX2_x, MVLX2_y).
2. The method of claim 1 , the method further comprising:
using different interpolation filters for motion compensation for the video block in different steps of the multi-step refinement process.
3. The method of claim 1 , wherein (MVLX2_x-MVLX1_x)<=Tx and (MVLX2_y-MVLX1_y)<=Ty, and wherein Tx and Ty represent thresholds respectively, and are predefined or signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a tile, a coding tree unit (CTU), or a coding unit (CU).
4. The method of claim 1 , wherein |MVLX2_x-MVLX1_x|<=Tx and |MVLX2_y-MVLX1_y|<=Ty wherein Tx and Ty represent thresholds respectively, and are predefined or signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a tile, a coding tree unit (CTU), or a coding unit (CU).
5. The method of claim 1 , wherein the refining comprises:
using first refined MVs (MVLX1_x, MVLX1_y), as a start searching point to derive the second refined MVs (MVLX2_x, MVLX2_y), or
using original MVs (MVLX0_x, MVLX0_y), which are not subjected to refinement, to derive the first refined MVs (MVLX1_x, MVLX1_y).
6. The method of claim 1 , wherein the multi-step refinement process is used in a decoder side motion vector refinement approach or a Bi-directional optical flow approach, and wherein the original MVs are signaled.
7. The method of claim 1 , further comprising:
modifying the first refined MVs of a first step of the multi-step refinement process before using the first refined MVs to generate at least one first motion compensated reference block or to derive the second refined MVs.
8. The method of claim 1 , wherein the video block corresponds to one of multiple sub-blocks, wherein a current block is split into the multiple sub-blocks before the multi-step refinement process is used,
wherein motion information of different sub-blocks from the multiple sub-blocks is refined with different numbers of steps of the multi-step refinement processing, and
wherein the refined MVs are derived at different sub-block sizes in different steps of the multi-step refinement process.
9. The method of claim 1 , wherein the video block corresponds to a current block, and the current block is split into multiple sub-blocks after a first step of the multi-step refinement process is used; and
motion information of at least one of the multiple sub-blocks is further refined.
10. The method of claim 1 , further comprising, in at least one step of the multi-step refinement process,
using the refined MVs to perform motion compensation for the video block; and
generating a prediction of the video block based on the motion compensation by using a Bi-directional optical flow approach, further comprising:
averaging respective predictions generated in different steps of the multi-step refinement process with weights to generate a final prediction of the video block.
11. The method of claim 1 , wherein the multi-step refinement processing is selectively used based on characteristics of the video block.
12. The method of claim 11 , wherein the characteristics of the video block comprises at least one of block size, coding mode information, motion information, position of the video block, and a type of slice, picture, or tile including the video block.
13. The method of claim 12 , wherein the multi-step refinement processing is not used if the video block contains luma samples whose number is less than a first threshold or more than a second threshold.
14. The method of claim 13 , wherein at least one of the first and second thresholds is selected from a group of 16, 32, or 64.
15. The method of claim 14 , wherein the multi-step refinement processing is not used when the video block has a minimum size of a width and height which is not larger than a third threshold, or
the multi-step refinement processing is not used when at least one of a width and height of the video block is not less than a fourth threshold, or
the multi-step refinement processing is not used when the video block has a size from a group of M×M, M×N, N×M, wherein M=128, N>=4 or N>=64, or
the multi-step refinement processing is not used when the video block is coded with advanced motion vector prediction (AMVP) mode in BIO approach or when the video block is coded with skip mode in BIO or DMVR approach.
16. The method of claim 1 , wherein the video block is a sub-block, wherein multi-step refinement processing depends on a position of the video block within a region covering the video block, and wherein the region corresponds to a prediction unit, a coding unit, a coding tree unit, a picture, or a tile.
17. The method of claim 1 , wherein the video processing comprises encoding the video block into a bitstream of the video block.
18. The method of claim 1 , wherein the video processing comprises decoding the video block from a bitstream of the video block.
19. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:
refine motion information of a video block by using a multi-step refinement processing, wherein multiple refined motion vectors (MVs) of the video block are derived iteratively in respective steps of the multi-step refinement processing, and
perform a video processing on the video block based on the multiple refined MVs of the video block;
wherein the refining comprises:
acquiring original MVs (MVLX0_x, MVLX0_y), to generate at least one original motion compensated reference block, wherein LX=L0 or L1, L0 and L1 representing reference list 0 and list 1 respectively,
refining, based on at least one original motion compensated reference block, the original MVs (MVLX0_x, MVLX0_y) to derive first refined MVs (MVLX1_x, MVLX1_y), to generating at least one first motion compensated reference block, and
refining the first refined MVs (MVLX1_x, MVLX1_y) to derive second refined MVs (MVLX2_x, MVLX2_y), wherein the first motion compensated reference block is used to derive at least one of temporal gradients, and spatial gradients used for deriving the second refined MVs (MVLX2_x, MVLX2_y).
20. A non-transitory computer readable medium storing instructions that cause a processor to:
refine motion information of a video block by using a multi-step refinement processing, wherein multiple refined motion vectors (MVs) of the video block are derived iteratively in respective steps of the multi-step refinement processing, and
perform a video processing on the video block based on the multiple refined MVs of the video block;
wherein the refining comprises:
acquiring original MVs (MVLX0_x, MVLX0_y), to generate at least one original motion compensated reference block, wherein LX=L0 or L1, L0 and L1 representing reference list 0 and list 1 respectively,
refining, based on at least one original motion compensated reference block, the original MVs (MVLX0_x, MVLX0_y) to derive first refined MVs (MVLX1_x, MVLX1_y), to generating at least one first motion compensated reference block, and
refining the first refined MVs (MVLX1_x, MVLX1_y) to derive second refined MVs (MVLX2_x, MVLX2_y), wherein the first motion compensated reference block is used to derive at least one of temporal gradients, and spatial gradients used for deriving the second refined MVs (MVLX2_x, MVLX2_y).