Systems and methods for joint coding of motion vector difference using template matching based scaling factor derivation
Systems and methods for joint coding of motion vector difference using template matching based scaling factor derivation include receiving a current frame in a video bitstream, determining that a current block in the current frame is coded in a joint motion vector difference (JMVD) mode, selecting first neighboring reconstructed samples and second neighboring reconstructed samples of the current block as template areas used for predicting the current block in the JVMD mode, determining a prediction block from a reference frame based on a scaling factor derived from the selected template areas and applied to a motion vector difference (MVD) associated with the current block, and reconstructing the current block in the JVMD mode based at least on the prediction block.
1 . A method for video decoding in a decoder, the method comprising:
receiving a current frame in a video bitstream comprising a current block and a plurality of neighboring blocks;
determining that the current block is coded in a joint motion vector difference (JMVD) mode;
in response to the current block being coded in the JMVD mode, selecting first neighboring reconstructed samples as a first template area, and second neighboring reconstructed samples of the current block as a second template area, used for predicting the current block in the JVMD mode;
determining a prediction block from a reference frame based on a scaling factor derived from the selected first and second template areas and applied to a motion vector difference (MVD) associated with the current block;
predicting a first motion vector prediction (MVP) that corresponds to a backward reference frame that is prior to the current frame in a display order;
predicting a second MVP that corresponds to a forward reference frame that is after the current frame in the display order; and
reconstructing the current block in the JVMD mode based at least on the prediction block.
2 . The method of claim 1 , wherein determining the prediction block from the reference frame based on the scaling factor comprises:
identifying a plurality of candidate scaling factors among a plurality of predetermined scaling factors; and
determining a scaling factor, among the plurality of candidate scaling factors, for the MVD associated with the current block, based on the first template area for the current block, the second template area for the current block, and the plurality of candidate scaling factors.
3 . The method of claim 1 , wherein the determining the scaling factor comprises:
generating a plurality of first prediction blocks, each of the plurality of first prediction blocks corresponding to a first candidate scaling factor among the plurality of candidate scaling factors; and
generating a plurality of second prediction blocks, each of the plurality of second prediction blocks corresponding to a second candidate scaling factor among the plurality of candidate scaling factors.
4 . The method of claim 3 , wherein generating the plurality of first prediction blocks comprises:
determining a plurality of first MVDs based on the first MVP and a plurality of first reference frames;
scaling each of the plurality of first MVDs based on the plurality of candidate scaling factors to obtain a plurality of scaled first MVDs for each of the plurality of scaling factors; and
generating a first prediction block, for each candidate scaling factor among the plurality of candidate scaling factors, using a motion vector equal to a sum of the first MVP and the plurality of scaled first MVDs corresponding to each candidate scaling factor.
5 . The method of claim 3 , wherein generating the plurality of second prediction blocks comprises:
determining a plurality of second MVDs based on the second MVP and a plurality of second reference frames;
scaling each of the plurality of second MVDs based on the plurality of candidate scaling factors to obtain a plurality of scaled second MVDs for each of the plurality of scaling factors; and
generating a second prediction block, for each candidate scaling factor among the plurality of candidate scaling factors, using a motion vector equal to a sum of the second MVP and the plurality of scaled second MVDs corresponding to each candidate scaling factor.
6 . The method of claim 3 , further comprising:
determining a first template area in each of a plurality of first reference frames, based on the plurality of first prediction blocks, the first template area in each first reference frame corresponding to the first template area for the current block in the current frame;
determining a second template area in each of the plurality of first reference frames, based on the plurality of first prediction blocks, that correspond to the second template area for the current block, the second template area in each first reference frame corresponding to the second template area for the current block in the current frame;
determining a first template area in each of a plurality of second reference frames, based on the plurality of second prediction blocks, the first template area in each second reference frame corresponding to the first template area for the current block in the current frame; and
determining a second template area in each of the plurality of second reference frames, based on the plurality of second prediction blocks, the second template area in each second reference frame corresponding to the second template area for the current block in the current frame.
7 . The method of claim 6 , further comprising:
generating a plurality of first template predictions associated with the current frame, based on a weighted average of the first template area associated with each first prediction block and the first template area associated with each second prediction block; and
generating a plurality of second template predictions associated with the current frame, based on a weighted average of the second template area associated with each first prediction block and the second template area associated with each second prediction block.
8 . The method of claim 7 , further comprising:
determining a first difference between each first template prediction and the first template area for the current frame;
determining a second difference between each second template prediction and the second template area for the current frame; and
evaluating the first and second differences based on a cost criterion, wherein a scaling factor associated with a lowest cost based on the cost criterion is determined as the scaling factor for the MVD associated with the current block.
9 . The method of claim 8 , wherein the cost criterion includes a sum of absolute difference (SAD), a sum of squared error (SSE), or a sum of absolute transform difference (SATD).
10 . A method of video encoding, the method comprising:
receiving a current frame comprising a current block and a plurality of neighboring blocks;
determining that the current block is to be coded in a joint motion vector difference (JMVD) mode; and
encoding the current block in the JMVD mode, encoding the current block in the JMVD mode signals to:
select, in response to the current block being coded in the JMVD mode, first neighboring reconstructed samples as a first template area, and second neighboring reconstructed samples of the current block as a second template area, used for predicting the current block in the JVMD mode;
determine a prediction block from a reference frame based on a scaling factor derived from the selected first and second template areas and applied to a motion vector difference (MVD) associated with the current block;
predict a first motion vector prediction (MVP) that corresponds to a backward reference frame that is prior to the current frame in a display order;
predict a second MVP that corresponds to a forward reference frame that is after the current frame in the display order; and
reconstruct the current block in the JVMD mode based at least on the prediction block.
11 . The method of claim 10 , wherein encoding the current block in the JMVD mode further signals to: generate a plurality of first prediction blocks, wherein each of the plurality of first prediction blocks corresponds to a first candidate scaling factor among the plurality of candidate scaling factors; and
generate a plurality of second prediction blocks, wherein each of the plurality of second prediction blocks corresponds to a second candidate scaling factor among the plurality of candidate scaling factors.
12 . The method of claim 11 , wherein encoding the current block in the JMVD mode further signals to: generate a plurality of first template predictions associated with the current frame, based on a weighted average of a first template area in each of a plurality of first reference frames, based on the plurality of first prediction blocks, and a first template area in each of a plurality of second reference frames, based on the plurality of second prediction blocks; and
generate a plurality of second template predictions associated with the current frame, based on a weighted average of a second template area in each first reference frame, based on the plurality of first prediction blocks, and a second template area in each second reference frame, based on the plurality of second prediction blocks.
13 . The method of claim 12 , wherein encoding the current block in the JMVD mode further signals to: determine a first difference between each first template prediction and the first template area for the current frame;
determine a second difference between each second template prediction and the second template area for the current frame; and
evaluate the first and second differences based on a cost criterion, wherein a scaling factor associated with a lowest cost based on the cost criterion is determined as the scaling factor for the MVD associated with the current block.
14 . A method of encoding visual media data, the method comprising:
generating a bitstream, comprising a current block and a plurality of neighboring blocks of the visual media data according to an encoding process including:
determining that the current block is to be coded in a joint motion vector difference (JMVD) mode; and
encoding the current block in the JMVD mode, encoding the current block in the JMVD mode signals to:
in response to the current block being coded in the JMVD mode, select first neighboring reconstructed samples as a first template area, and second neighboring reconstructed samples of the current block as a second template area, used for predicting the current block in the JVMD mode;
determine a prediction block from a reference frame based on a scaling factor derived from the selected first and second template areas and applied to a motion vector difference (MVD) associated with the current block;
predict a first motion vector prediction (MVP) that corresponds to a backward reference frame that is prior to a current frame in a display order;
predict a second MVP that corresponds to a forward reference frame that is after the current frame in the display order; and
reconstruct the current block in the JVMD mode based at least on the prediction block; and
transmitting the generated bitstream of the visual media data.
15 . The non-transitory computer readable medium of claim 14 , wherein encoding the current block in the JMVD mode further signals to:
generate a plurality of first prediction blocks, wherein each of the plurality of first prediction blocks corresponds to a first candidate scaling factor among the plurality of candidate scaling factors; and
generate a plurality of second prediction blocks, wherein each of the plurality of second prediction blocks corresponds to a second candidate scaling factor among the plurality of candidate scaling factors.
16 . The non-transitory computer readable medium of claim 15 , wherein encoding the current block in the JMVD mode further signals to:
generate a plurality of first template predictions associated with the current frame, based on a weighted average of a first template area in each of a plurality of first reference frames, based on the plurality of first prediction blocks, and a first template area in each of a plurality of second reference frames, based on the plurality of second prediction blocks; and
generate a plurality of second template predictions associated with the current frame, based on a weighted average of a second template area in each first reference frame, based on the plurality of first prediction blocks, and a second template area in each second reference frame, based on the plurality of second prediction blocks.
17 . The non-transitory computer readable medium of claim 16 , wherein encoding the current block in the JMVD mode further signals to:
determine a first difference between each first template prediction and the first template area for the current frame;
determine a second difference between each second template prediction and the second template area for the current frame; and
evaluate the first and second differences based on a cost criterion, wherein a scaling factor associated with a lowest cost based on the cost criterion is determined as the scaling factor for the MVD associated with the current block.