Context derivation for motion vector difference coding
This disclosure relates encoding and decoding of motion vector difference in for inter-predicting a video block. An example is disclosed for decoding an inter-predicted video block of a video stream. The method includes deriving an inter-prediction mode for the video block from the video stream; deriving a motion vector prediction mode for the video block; deriving, from the video stream, a context for signaling a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the inter-prediction mode and/or the motion vector prediction mode; and decoding the video block based on the set of syntax elements and the derived context.
1. A method for decoding a video block in a video stream, the method comprising:
receiving the video stream;
determining an inter-prediction mode for the video block from the video stream;
determining a motion vector prediction mode for the video block, wherein the motion vector prediction mode comprises at least one of: a NEARMV mode, a NEWMV mode, a NEW_NEARMV mode, a GLOBALMV mode, a NEAR_NEWMV mode, a NEW_NEWMV mode, or a GOLBAL GLOBALMV mode;
deriving, from the video stream, a context for signaling a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the motion vector prediction mode, wherein the set of syntax elements associated with the MVD comprises at least one of:
an mv_joint syntax element indicating non-zero components of the MVD;
an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;
an mv_class syntax element indicating an MVD class of the MVD;
an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;
an mv_fr syntax element indicating first two fractional bits of the MVD; and
an mv_hp syntax element indicating a third fractional bit of the MVD; and
decoding the video block based on the set of syntax elements and the derived context.
2. The method of claim 1 , wherein the video block comprises one of: a coded block, a prediction block, or a coding unit.
3. The method of claim 1 , wherein the inter-prediction mode indicates whether motion vector difference or differences are signaled for a single inter-prediction reference frame list or signaled for both of two inter-prediction reference frame lists.
4. The method of claim 3 , wherein deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:
in response to the MVD being signaled for the single inter-prediction reference frame list, deriving the context as a first predefined context; and
in response to the MVD being signaled for both of the two inter-prediction reference frame lists, deriving the context as a second predefined context,
the first predefined context and the second predefined context being different.
5. The method of claim 1 , wherein the inter-prediction mode comprising one of a single-reference inter-prediction mode or a compound-reference inter-prediction mode.
6. The method of claim 5 , wherein deriving the context comprises deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block based on at least the inter-prediction mode.
7. The method of claim 6 , wherein deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:
in response to the inter-prediction mode being the single-reference inter-prediction mode, deriving the context as a first predefined context; and
in response to the inter-prediction mode being the compound-reference inter-prediction mode, deriving the context as a second predefined context,
the first predefined context and the second predefined context being different.
8. The method of claim 5 , wherein:
for the single-reference inter-prediction mode, the motion vector prediction mode comprises one of a direct merge motion vector predictor mode (NEARMV), a merge mode motion vector difference prediction mode (NEWMV), and a global motion vector predictor mode (GLOBALMV) associated with a single inter-prediction reference frame for the video block; and
for the compound-reference inter-prediction mode with two reference inter-prediction frames, the motion vector prediction mode comprises one of:
the direct merge motion vector predictor mode for the first of the two reference inter-prediction frames and the merge mode motion vector difference prediction mode for the second of the two reference inter-prediction frames (NEAR_NEWMV);
the direct merge motion vector predictor mode for the second of the two reference inter-prediction frames and the merge mode motion vector difference prediction mode for the first of the two reference inter-prediction frames (NEW_NEARMV);
the merge mode motion vector difference prediction mode for both the first and the second of the two reference inter-prediction frames (NEW_NEWMV); and
the global motion vector predictor mode for both the first and the second of the two reference inter-prediction frames (GLOBAL_GLOBALMV).
9. The method of claim 8 , wherein deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:
in response to the motion vector prediction mode being NEWMV or NEW_NEWMV, deriving the context as a first predefined context; and
in response to the motion vector prediction mode not being NEWMV or NEW_NEWMV, deriving the context as a second predefined context,
the first predefined context and the second predefined context being different.
10. The method of claim 9 , wherein the first predefined context and the second predefined context are configured for signaling one of: an mv_joint syntax for indicating non-zero components of the MVD.
11. The method of claim 8 , wherein the inter-prediction mode comprises the compound-reference inter-prediction mode, and wherein deriving the context comprises deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block based on at least on whether the motion vector prediction mode is one of NEW_NEARMV or NEAR_NEWMV.
12. The method of claim 11 , wherein deriving the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:
in response to the motion vector prediction mode being one of NEW_NEARMV or the NEAR_NEWMV, deriving the context as a first predefined context; and
in response to the motion vector prediction mode not being one of NEW_NEARMV or the NEAR_NEWMV, deriving the context as a second predefined context,
the first predefined context being different from the second predefined context.
13. The method of claim 12 , wherein the first predefined context and the second predefined context are configured for signaling an mv_joint syntax for indicating non-zero MVD components.
14. The method of claim 12 , wherein the first predefined context and the second predefined context are configured for signaling an mv_class syntax indicating an MVD class of the MVD.
15. Device for decoding a video block in a video stream, the device comprising a memory for storing computer instructions and a processor in communication with the memory, wherein, when the processor executes the computer instructions, the processor is configured to cause the device to:
receive a video stream;
determine an inter-prediction mode for the video block from the video stream;
determine a motion vector prediction mode for the video block, wherein the motion vector prediction mode comprises at least one of: a NEARMV mode, a NEWMV mode, a NEW_NEARMV mode, a GLOBALMV mode, a NEAR_NEWMV mode, a NEW_NEWMV mode, or a GOLBAL GLOBALMV mode;
derive, from the video stream, a context for signaling a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the motion vector prediction mode, wherein the set of syntax elements associated with the MVD comprises at least one of:
an mv_joint syntax element indicating non-zero components of the MVD;
an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;
an mv_class syntax element indicating an MVD class of the MVD;
an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;
an mv_fr syntax element indicating first two fractional bits of the MVD; and
an mv_hp syntax element indicating a third fractional bit of the MVD; and
decode the video block based on the set of syntax elements and the derived context.
16. The device of claim 15 , wherein:
the inter-prediction mode indicates whether motion vector difference or differences are signaled for a single inter-prediction reference frame list or signaled for both of two inter-prediction reference frame lists; and
when the processor is configured to cause the device to derive the context for signaling the set of syntax elements associated with the MVD associated with the video block, the processor is configured to cause the device to:
in response to the MVD being signaled for the single inter-prediction reference frame list, derive the context as a first predefined context; and
in response to the MVD being signaled for both of the two inter-prediction reference frame lists, derive the context as a second predefined context,
the first predefined context and the second predefined context being different.
17. The device of claim 16 , wherein:
the inter-prediction mode comprising one of a single-reference inter-prediction mode or a compound-reference inter-prediction mode; and
when the processor is configured to cause the device to derive the context, the processor is configured to cause the device to derive the context for signaling the set of syntax elements associated with the MVD associated with the video block based on at least the inter-prediction mode.
18. The device of claim 17 , wherein, when the processor is configured to cause the device to derive the context for signaling the set of syntax elements associated with the MVD associated with the video block, the processor is configured to cause the device to:
in response to the inter-prediction mode being the single-reference inter-prediction mode, derive the context as a first predefined context; and
in response to the inter-prediction mode being the compound-reference inter-prediction mode, derive the context as a second predefined context,
the first predefined context and the second predefined context being different.
19. A non-transitory storage medium for storing computer readable instructions, the computer readable instructions, when executed a processor, causing the processor to:
receive a coded video stream;
determine an inter-prediction mode for a video block from the coded video stream;
determine a motion vector prediction mode for the video block;
derive, from the video stream, a context for signaling a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the motion vector prediction mode, wherein the set of syntax elements associated with the MVD comprises at least one of:
an mv_joint syntax element indicating non-zero components of the MVD;
an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;
an mv_class syntax element indicating an MVD class of the MVD;
an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;
an mv_fr syntax element indicating first two fractional bits of the MVD; and
an mv_hp syntax element indicating a third fractional bit of the MVD; and
decode the video block based on the set of syntax elements and the derived context.