Motion vector prediction for video coding
A method of decoding a video signal, apparatus, and a non-transitory computer-readable storage medium are provided. The method includes obtaining a video block from the video signal, obtaining spatial neighboring blocks based on the video block, obtaining up to one left non-scaled motion vector predictor (MVP) based on the multiple left spatial neighboring blocks, obtaining up to one above non-scaled MVP based on the multiple above spatial neighboring blocks, deriving an MVP candidate list based on the coding block, the multiple left spatial neighboring blocks, and the multiple above spatial neighboring blocks by performing no motion vector scaling in all scenarios for deriving motion vectors in deriving process for spatial MVP candidates, receiving an index indicating an MVP in the MVP candidate list, and obtaining a prediction signal of the coding block based on the MVP indicated by the index.
1 . A method of video decoding, comprising:
obtaining a coding block from a video signal;
obtaining a plurality of spatial neighboring blocks based on the coding block, wherein the plurality of spatial neighboring blocks comprise multiple left spatial neighboring blocks and multiple above spatial neighboring blocks;
determining up to one left non-scaled spatial motion vector predictor (MVP) candidate based on the multiple left spatial neighboring blocks, and determining up to one above non-scaled spatial MVP candidate based on the multiple above spatial neighboring blocks;
deriving an MVP candidate list based on the up to one left non-scaled spatial MVP candidate, and the up to one above non-scaled spatial MVP candidate;
receiving an index indicating an MVP in the MVP candidate list; and
obtaining a prediction signal of the coding block based on the MVP indicated by the index,
wherein determining the up to one left non-scaled spatial MVP candidate and the up to one above non-scaled spatial MVP candidate comprises:
determining whether a picture order count (POC) of a reference picture of one of the plurality of spatial neighboring blocks is different from a POC of a reference picture of the coding block; and
in a case that the POC of the reference picture of the one of the plurality of spatial neighboring blocks is different from the POC of the reference picture of the coding block, not performing operations for scaling a motion vector.
2 . The method of claim 1 , further comprising:
determining up to one temporal MVP candidate from temporal collocated blocks,
wherein deriving an MVP candidate list based on the up to one left non-scaled spatial MVP candidate, and the up to one above non-scaled spatial MVP candidate comprises:
deriving the MVP candidate list based on the up to one left non-scaled spatial MVP candidate, the up to one above non-scaled spatial MVP candidate, and the up to one temporal MVP candidate.
3 . The method of claim 2 , further comprising:
obtaining an explicit flag indicating whether a collocated picture used for temporal motion vector prediction is derived from a first reference picture list or a second reference picture list.
4 . The method of claim 3 , further comprising:
obtaining a collocated reference index specifying a reference index of the collocated picture used for temporal motion vector prediction.
5 . The method of claim 2 , wherein determining up to one temporal MVP candidate from temporal collocated blocks comprises:
performing motion vector scaling on an MVP derived from the temporal collocated blocks.
6 . The method of claim 5 , wherein the motion vector scaling is performed based on a picture order count (POC) difference between a reference picture of a current picture and the current picture, and a POC difference between a reference picture of a collocated picture and the collocated picture.
7 . The method of claim 2 , wherein deriving the MVP candidate list comprises:
pruning, when there are two identical MVP candidates derived from the plurality of spatial neighboring blocks, one of the two identical MVP candidates from the MVP candidate list;
determining up to two history-based MVP candidates from a First In First Out (FIFO) table; and
obtaining up to two MVs equal to 0,
wherein the temporal collocated blocks comprise a bottom-right collocated block T0 and a central collocated block T1.
8 . The method of claim 1 , wherein the plurality of spatial neighboring blocks comprise a below left block, a left block, an above right block, an above block, and an above left block.
9 . A computing device comprising:
one or more processors;
a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors are configured to:
obtain a coding block from a video signal;
obtain a plurality of spatial neighboring blocks based on the coding block, wherein the plurality of spatial neighboring blocks comprise multiple left spatial neighboring blocks and multiple above spatial neighboring blocks;
determine up to one left non-scaled spatial motion vector predictor (MVP) candidate based on the multiple left spatial neighboring blocks, and determine up to one above non-scaled spatial MVP candidate based on the multiple above spatial neighboring blocks;
derive an MVP candidate list based on the up to one left non-scaled spatial MVP candidate, and the up to one above non-scaled spatial MVP candidate;
determine an index indicating an MVP in the MVP candidate list; and
obtain a prediction signal of the coding block based on the MVP indicated by the index,
wherein determining the up to one left non-scaled spatial MVP candidate and the up to one above non-scaled spatial MVP candidate comprises:
determining whether a picture order count (POC) of a reference picture of one of the plurality of spatial neighboring blocks is different from a POC of a reference picture of the coding block; and
in a case that the POC of the reference picture of the one of the plurality of spatial neighboring blocks is different from the POC of the reference picture of the coding block, not performing operations for scaling a motion vector.
10 . The computing device of claim 9 , wherein the one or more processors are configured to derive the MVP candidate list by:
pruning, when there are two identical MVP candidates derived from the plurality of spatial neighboring blocks, one of the two identical MVP candidates from the MVP candidate list;
obtaining up to one MVP from temporal collocated blocks, wherein the temporal collocated blocks comprise a bottom-right collocated block T0 and a central collocated block T1;
determining up to two history-based MVP candidates from a First In First Out (FIFO) table; and
obtaining up to two MVs equal to 0.
11 . The computing device of claim 10 , wherein the one or more processors are further configured to:
obtain an explicit flag indicating whether a collocated picture used for temporal motion vector prediction is derived from a first reference picture list or a second reference picture list.
12 . The computing device of claim 11 , wherein the one or more processors are further configured to:
obtain a collocated reference index specifying a reference index of the collocated picture used for temporal motion vector prediction.
13 . The computing device of claim 10 , wherein the one or more processors are configured to obtain up to one MVP from temporal collocated blocks by:
performing motion vector scaling on an MVP derived from the temporal collocated blocks.
14 . The computing device of claim 13 , wherein the motion vector scaling is performed based on a picture order count (POC) difference between a reference picture of a current picture and the current picture, and a POC difference between a reference picture of a collocated picture and the collocated picture.
15 . The computing device of claim 9 , wherein the plurality of spatial neighboring blocks comprise a below left block, a left block, an above right block, an above block, and an above left block.
16 . A method for storing a bitstream, comprising:
generating the bitstream by performing an encoding method; and
storing the bitstream in a non-transitory computer-readable storage medium, wherein the encoding method comprises:
determining a coding block for a video signal;
determining a plurality of spatial neighboring blocks based on the coding block, wherein the plurality of spatial neighboring blocks comprise multiple left spatial neighboring blocks and multiple above spatial neighboring blocks;
determining up to one left non-scaled spatial motion vector predictor (MVP) candidate based on the multiple left spatial neighboring blocks, and determining up to one above non-scaled spatial MVP candidate based on the multiple above spatial neighboring blocks;
deriving an MVP candidate list based on the up to one left non-scaled spatial MVP candidate, and the up to one above non-scaled spatial MVP candidate; and
signaling an index indicating an MVP to be used for prediction of the coding block in the MVP candidate list,
wherein determining the up to one left non-scaled spatial MVP candidate and the up to one above non-scaled spatial MVP candidate comprises:
determining whether a picture order count (POC) of a reference picture of one of the plurality of spatial neighboring blocks is different from a POC of a reference picture of the coding block; and
in a case that the POC of the reference picture of the one of the plurality of spatial neighboring blocks is different from the POC of the reference picture of the coding block, not performing operations for scaling a motion vector.
17 . The method of claim 16 , wherein the encoding method further comprises:
pruning, when there are two identical MVP candidates derived from the plurality of spatial neighboring blocks, one of the two identical MVP candidates from the MVP candidate list;
determining up to one MVP from temporal collocated blocks, wherein the temporal collocated blocks comprise a bottom-right collocated block T0 and a central collocated block T1;
determining up to two history-based MVP candidates from a First In First Out (FIFO) table; and
determining up to two MVs equal to 0.
18 . The method of claim 17 , wherein the encoding method further comprises:
signaling an explicit flag indicating whether a collocated picture used for temporal motion vector prediction is derived from a first reference picture list or a second reference picture list.
19 . The method of claim 18 , wherein the encoding method further comprises:
signaling a collocated reference index specifying a reference index of the collocated picture used for temporal motion vector prediction.
20 . The method of claim 16 , wherein the plurality of spatial neighboring blocks comprise a below left block, a left block, an above right block, an above block, and an above left block.