Subblock based motion vector predictor displacement vector reordering using template matching
Aspects of the disclosure provide a method and an apparatus for video encoding/decoding. The apparatus includes processing circuitry for: receiving prediction information of a current coding block in a current picture from a coded video bitstream, the prediction information indicating that the current coding block is coded using a subblock-based temporal motion vector prediction (SbTMVP) mode; deriving multiple displacement vector (DV) candidates by applying multiple DV offset candidates to a fixed DV predictor of the current coding block; comparing a template of the current coding block with each of multiple templates, each template of the multiple templates being located at a position specified by a corresponding one of the multiple DV candidates; calculating a cost value associated with each one of the multiple DV offset candidates based on the comparing; and reordering DV offset indices of the multiple DV offset candidates based on their calculated cost values.
1 . A method of video decoding, the method comprising:
receiving prediction information of a current coding block in a current picture from a coded video bitstream, the prediction information indicating that the current coding block is coded using a subblock-based temporal motion vector prediction (SbTMVP) mode;
comparing a template of the current coding block with each of multiple templates, each template of the multiple templates being located at a position specified by a corresponding one of multiple displacement vector (DV) candidates;
calculating a cost value associated with each one of the multiple DV candidates based on the comparison of the template with each of the multiple templates;
selecting a DV candidate from the multiple DV candidates with a lowest calculated cost value or from a reordered list of the multiple DV candidates based on the calculated cost values; and
predicting the current coding block in the SbTMVP mode based at least on the selected DV candidate.
2 . The method of claim 1 , further comprising:
obtaining multiple DV offsets from the coded video bitstream, each DV offset corresponding to a respective one of the multiple DV candidates; and
deriving the multiple DV candidates by applying the multiple DV offsets to a fixed DV predictor of the current coding block.
3 . The method of claim 1 , wherein the predicting comprises:
predicting the current coding block in the SbTMVP mode based at least on an index of the reordered list of the multiple DV candidates that is signaled in the coded video bitstream, the index indicating which DV candidate is selected from the reordered list of the multiple DV candidates for performing SbTMVP.
4 . The method of claim 1 , wherein the selecting the DV candidate from the reordered list comprises selecting the DV candidate of the multiple DV candidates with a lowest calculated cost value by default for performing SbTMVP.
5 . The method of claim 1 , wherein the cost value is calculated by performing Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Squared Error (SSE), sub-sampled SAD, or mean-removed SAD.
6 . The method of claim 1 , wherein the multiple DV candidates comprise Merge with Motion Vector Difference (MMVD) candidates.
7 . The method of claim 6 , wherein the comparing, the calculating, and the selecting are performed only for a subset of the MMVD candidates, wherein a relative order of one or more other ones of the MMVD candidates is kept unchanged.
8 . The method of claim 6 , wherein the comparing, the calculating, and the selecting are performed for all of the MMVD candidates, wherein after reordering only a number N of the MMVD candidates which have lowest cost values are used, wherein the number N is less than or equal to a total number of the MMVD candidates.
9 . The method of claim 1 , wherein the multiple DV candidates include multiple DV predictor candidates.
10 . The method of claim 9 , further comprising:
receiving an index signaled in the coded video bitstream, wherein the index indicates which DV predictor candidate is selected from the reordered list of the multiple DV candidates for performing SbTMVP.
11 . The method of claim 10 , wherein the method further comprises selecting a DV predictor candidate with a lowest calculated cost value by default for performing SbTMVP.
12 . The method of claim 10 , wherein the list of the multiple DV candidates is constructed from spatial neighboring coding units (CUs) or from history-based motion vector prediction (HMVP) candidates.
13 . The method of claim 10 , wherein only a first N number of the multiple DV predictor candidates on the reordered list of the multiple DV candidates are signaled.
14 . The method of claim 1 , wherein the list of the multiple DV candidates includes at least (i) a first DV candidate that is derived by applying a DV offset to a fixed DV predictor of the current coding block and (ii) a second DV candidate that includes a DV predictor candidate.
15 . A method of video encoding, the method comprising:
determining that a current coding block in a current picture is to be coded using a subblock-based temporal motion vector prediction (SbTMVP) mode;
comparing a template of the current coding block with each of multiple templates, each template of the multiple templates being located at a position specified by a corresponding one of multiple displacement vector (DV) candidates;
calculating a cost value associated with each one of the multiple DV candidates based on the comparison of the template with each of the multiple templates;
selecting a DV candidate from the multiple DV candidates with a lowest calculated cost value or from a reordered list of the multiple DV candidates based on the calculated cost values; and
encoding the current coding block in the SbTMVP mode in a bitstream based at least on the selected DV candidate.
16 . The method of claim 15 , further comprising:
determining multiple DV offsets, each DV offset corresponding to a respective one of the multiple DV candidates; and
deriving the multiple DV candidates by applying the multiple DV offsets to a fixed DV predictor of the current coding block.
17 . The method of claim 15 , wherein the selecting the DV candidate from the reordered list comprises selecting the DV candidate of the multiple DV candidates with a lowest calculated cost value by default for performing SbTMVP.
18 . A non-transitory computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to perform a method of encoding a bitstream comprising:
determining that a current coding block in a current picture is to be coded using a subblock-based temporal motion vector prediction (SbTMVP) mode;
comparing a template of the current coding block with each of multiple templates, each template of the multiple templates being located at a position specified by a corresponding one of multiple displacement vector (DV) candidates;
calculating a cost value associated with each one of the multiple DV candidates based on the comparison of the template with each of the multiple templates;
selecting a DV candidate from the multiple DV candidates with a lowest calculated cost value or from a reordered list of the multiple DV candidates based on the calculated cost values;
encoding the current coding block in the SbTMVP mode in the bitstream based at least on the selected DV candidate; and
transmitting the encoded bitstream.
19 . The non-transitory computer-readable storage medium of claim 18 , wherein the method further comprises:
determining multiple DV offsets, each DV offset corresponding to a respective one of the multiple DV candidates; and
deriving the multiple DV candidates by applying the multiple DV offsets to a fixed DV predictor of the current coding block.
20 . The non-transitory computer-readable storage medium of claim 18 , wherein the selecting the DV candidate from the reordered list comprises selecting the DV candidate of the multiple DV candidates with a lowest calculated cost value by default for performing SbTMVP.