SYSTEMS AND METHODS FOR MOTION INFORMATION TRANSFER FROM VISUAL TO FEATURE DOMAIN AND FEATURE-BASED DECODER-SIDE MOTION VECTOR REFINEMENT CONTROL
Systems and methods for motion information transfer from visual to feature domain are disclosed which provide for mapping motion information from coding units in video content to corresponding convolution unit(s) in feature content. Systems and methods are also provided for improved decoder-side motion vector refinement of video content based on characteristics of corresponding feature units.
1 . A method of encoding video content with feature information, comprising:
determine motion information for each coding unit comprising the video content;
produce a feature map for the video content, the feature map having a plurality of convolution units with a correspondence to the convolution units;
using a transformation selected based on the correspondence of the convolution units to the coding units, generate motion transformation information by mapping the motion information of the video content in each coding unit to at least one corresponding convolution unit; and
generate an encoded bitstream including the video content, the motion information, and the motion transformation information.
2 . The method of claim 1 , wherein the coding units correspond to the convolution units in size and number and wherein the transformation copies the motion information from each coding unit to a corresponding convolution unit.
3 . The method of claim 1 , wherein the coding units correspond to the convolution units in number but differ in size and wherein the transformation scales the motion information from each coding unit to a corresponding convolution unit.
4 . The method of claim 1 , wherein multiple coding units correspond to a single convolution unit and wherein the transformation fuses the motion information of the multiple coding units to map the motion to the convolution unit.
5 . The method of claim 1 , wherein each coding unit corresponds to multiple convolution units and wherein the transformation merges the motion information from a coding unit to the multiple convolution units.
6 . The method of claim 1 , wherein the transformation is selected from a group including copying, scaling, fusing and merging.
7 . The method of claim 1 , wherein the bitstream comprises:
header and metadata for the entire content;
a video sub-bitstream including header, metadata, video payload information including motion information; and
a feature sub-bitstream including header, metadata and feature payload information including motion transform information.
8 . A method for decoding a bitstream, the bitstream having a video content comprising a plurality of coding units having associated motion vectors and feature content comprising a plurality of feature units encoded in the bitstream, the encoder having a mode for decoder-side motion vector refinement (DMVR), the method comprising:
for each feature unit, determine if the feature unit includes an object of interest;
for a coding unit corresponding to the feature unit, determine if the DMVR mode is enabled;
if the feature unit includes an object of interest and the DMVR mode for the corresponding coding unit is not enabled, enable the DMVR mode for that coding unit; and
if the feature unit does not include an object of interest and the DMVR mode for the corresponding coding unit is enabled, disable the DMVR mode for that coding unit.
9 . The method of claim 1 , wherein the status of the DMVR mode for each coding unit is signaled in the bitstream.
10 . The method of claim 9 , wherein the DMVR mode is signaled in a picture header of the bitstream.
11 . The method of claim 10 , wherein the DMVR mode is signaled in a sequence parameter set of the bitstream.
12 . A decoder for decoding a bitstream, the bitstream having a video content comprising a plurality of coding units having associated motion vectors and feature content comprising a plurality of feature units encoded in the bitstream, the encoder having a mode for decoder-side motion vector refinement (DMVR), the decoder comprising:
a video input receiving the bitstream; and
a processor programmed with instructions comprising:
for each feature unit, determine if the feature unit includes an object of interest;
for a coding unit corresponding to the feature unit, determine if the DMVR mode is enabled;
if the feature unit includes an object of interest and the DMVR mode for the corresponding coding unit is not enabled, enable the DMVR mode for that coding unit; and
if the feature unit does not include an object of interest and the DMVR mode for the corresponding coding unit is enabled, disable the DMVR mode for that coding unit.
13 . The decoder of claim 12 , wherein the status of the DMVR mode for each coding unit is signaled in the bitstream.
14 . The decoder of claim 13 , wherein the DMVR mode is signaled in a picture header of the bitstream.
15 . The decoder of claim 13 , wherein the DMVR mode is signaled in a sequence parameter set of the bitstream.