IP Library › Granted Patent US 12,641,244
Granted Patent B2
US 12,641,244 · App. 18/808,343 · Granted May 26, 2026

Context derivation for motion vector difference coding

Inventors: Liang Zhao (Sunnyvale, CA); Xin Zhao (San Jose, CA); Shan Liu (San Jose, CA)
Assignee: Tencent America LLC
H04N19/137H04N19/176H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,641,244
App. No.
18/808,343
Granted
May 26, 2026
Kind
B2
Abstract

This disclosure relates encoding and decoding of motion vector difference in for inter-predicting a video block. An example is disclosed for decoding an inter-predicted video block of a video stream. The method includes deriving an inter-prediction mode for the video block from the video stream; deriving a motion vector prediction mode for the video block; deriving, from the video stream, a context for signaling a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the inter-prediction mode and/or the motion vector prediction mode; and decoding the video block based on the set of syntax elements and the derived context.

Claims (66)

1 . A method for encoding a video block in a video stream, the method comprising:

determining a motion vector prediction mode for the video block, wherein the motion vector prediction mode comprises at least one of: a NEARMV mode, a NEWMV mode, a NEW_NEARMV mode, a GLOBALMV mode, a NEAR_NEWMV mode, a NEW_NEWMV mode, or a GOLBAL_GLOBALMV mode;

determining a context for encoding a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the motion vector prediction mode, wherein the set of syntax elements associated with the MVD comprises at least one of:

an mv_joint syntax element indicating non-zero components of the MVD;

an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;

an mv_class syntax element indicating an MVD class of the MVD;

an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;

an mv_fr syntax element indicating first two fractional bits of the MVD; and

an mv_hp syntax element indicating a third fractional bit of the MVD; and

encoding the video block based on the set of syntax elements and the determined context.

2 . The method of claim 1 , wherein the video block comprises one of: a coded block, a prediction block, or a coding unit.

3 . The method of claim 1 , wherein the motion vector prediction mode indicates whether motion vector difference or differences are signaled for a single inter-prediction reference frame list or signaled for both of two inter-prediction reference frame lists.

4 . The method of claim 3 , wherein determining the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:

in response to the MVD being signaled for the single inter-prediction reference frame list, determining the context as a first predefined context; and

in response to the MVD being signaled for both of the two inter-prediction reference frame lists, determining the context as a second predefined context,

the first predefined context and the second predefined context being different.

5 . The method of claim 1 , wherein the motion vector prediction mode comprising a single-reference inter-prediction mode or a compound-reference inter-prediction mode.

6 . The method of claim 5 , wherein determining the context comprises determining the context for signaling the set of syntax elements associated with the MVD associated with the video block based on at least the motion vector prediction mode.

7 . The method of claim 6 , wherein determining the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:

in response to the motion vector prediction mode being the single-reference inter-prediction mode, determining the context as a first predefined context; and

in response to the motion vector prediction mode being the compound-reference inter-prediction mode, determining the context as a second predefined context,

the first predefined context and the second predefined context being different.

8 . The method of claim 5 , wherein:

for the single-reference inter-prediction mode, the motion vector prediction mode comprises one of a direct merge motion vector predictor mode (NEARMV), a merge mode motion vector difference prediction mode (NEWMV), and a global motion vector predictor mode (GLOBALMV) associated with a single inter-prediction reference frame for the video block; and

for the compound-reference inter-prediction mode with two reference inter-prediction frames, the motion vector prediction mode comprises one of:

the direct merge motion vector predictor mode for a first of the two reference inter-prediction frames and the merge mode motion vector difference prediction mode for a second of the two reference inter-prediction frames (NEAR_NEWMV);

the direct merge motion vector predictor mode for the second of the two reference inter-prediction frames and the merge mode motion vector difference prediction mode for the first of the two reference inter-prediction frames (NEW_NEARMV);

the merge mode motion vector difference prediction mode for both the first and the second of the two reference inter-prediction frames (NEW_NEWMV); and

the global motion vector predictor mode for both the first and the second of the two reference inter-prediction frames (GLOBAL_GLOBALMV).

9 . The method of claim 8 , wherein determining the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:

in response to the motion vector prediction mode being the NEWMV mode or the NEW_NEWMV mode, determining the context as a first predefined context; and

in response to the motion vector prediction mode not being the NEWMV mode or the NEW_NEWMV mode, determining the context as a second predefined context,

the first predefined context and the second predefined context being different.

10 . The method of claim 9 , wherein the first predefined context and the second predefined context are configured for signaling one of: an mv_joint syntax for indicating non-zero components of the MVD.

11 . The method of claim 8 , wherein the motion vector prediction mode comprises the compound-reference inter-prediction mode, and wherein determining the context comprises determining the context for signaling the set of syntax elements associated with the MVD associated with the video block based on at least on whether the motion vector prediction mode is one of the NEW_NEARMV mode or the NEAR_NEWMV mode.

12 . The method of claim 11 , wherein determining the context for signaling the set of syntax elements associated with the MVD associated with the video block comprises:

in response to the motion vector prediction mode being one of the NEW_NEARMV mode or the NEAR_NEWMV mode, determining the context as a first predefined context; and

in response to the motion vector prediction mode not being one of the NEW_NEARMV or the NEAR_NEWMV, determining the context as a second predefined context,

the first predefined context being different from the second predefined context.

13 . The method of claim 12 , wherein the first predefined context and the second predefined context are configured for signaling an mv_joint syntax for indicating non-zero MVD components.

14 . The method of claim 12 , wherein the first predefined context and the second predefined context are configured for signaling an mv_class syntax indicating an MVD class of the MVD.

15 . A method for processing visual media file, comprising:

performing a conversion between a visual media file and a bitstream of a visual media data, wherein:

the bitstream comprises, for a video block, a motion vector prediction mode and a set of syntax elements associated with a Motion Vector Difference (MVD), the set of syntax elements being encoded with a context based on the motion vector prediction mode,

wherein the set of syntax elements associated with the MVD comprises at least one of:

an mv_joint syntax element indicating non-zero components of the MVD;

an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;

an mv_class syntax element indicating an MVD class of the MVD;

an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;

an mv_fr syntax element indicating first two fractional bits of the MVD; and

an mv_hp syntax element indicating a third fractional bit of the MVD; and

wherein the motion vector prediction mode comprises at least one of: a NEARMV mode, a NEWMV mode, a NEW_NEARMV mode, a GLOBALMV mode, a NEAR_NEWMV mode, a NEW_NEWMV mode, or a GOLBAL_GLOBALMV mode.

16 . The method of claim 15 , wherein the motion vector prediction mode indicates whether motion vector difference or differences are signaled for a single inter-prediction reference frame list or signaled for both of two inter-prediction reference frame lists.

17 . The method of claim 16 , wherein, when the context for encoding the set of syntax elements differs between when the MVD is signaled for a single inter-prediction reference frame list and for both a first inter-prediction reference frame and a second inter-prediction reference frame list.

18 . The method of claim 15 , wherein the motion vector prediction mode comprising a compound-reference inter-prediction mode.

19 . A device for encoding a video block in a video stream, the device comprising a memory for storing computer instructions and a processor in communication with the memory, wherein, when the processor executes the computer instructions, the processor is configured to cause the device to:

determine a motion vector prediction mode for the video block, wherein the motion vector prediction mode comprises at least one of: a NEARMV mode, a NEWMV mode, a NEW_NEARMV mode, a GLOBALMV mode, a NEAR_NEWMV mode, a NEW_NEWMV mode, or a GOLBAL_GLOBALMV mode;

determine a context for encoding a set of syntax elements associated with a Motion Vector Difference (MVD) associated with the video block based on the motion vector prediction mode, wherein the set of syntax elements associated with the MVD comprises at least one of:

an mv_joint syntax element indicating non-zero components of the MVD;

an mv_sign syntax element, which by itself alone specifies whether the MVD is positive or negative, without using other information;

an mv_class syntax element indicating an MVD class of the MVD;

an mv_bit syntax element indicating a corresponding integer part of an offset between the MVD and a starting magnitude of the MVD class;

an mv_fr syntax element indicating first two fractional bits of the MVD; and

an mv_hp syntax element indicating a third fractional bit of the MVD; and

encode the video block based on the set of syntax elements and the determined context.

20 . The device of claim 19 , wherein, when the processor is configured to cause the device to determine the context, the processor is configured to cause the device to determine the context for encoding the set of syntax elements associated with the MVD associated with the video block based on at least the motion vector prediction mode.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2026
From: ZHAO, LIANG; ZHAO, XIN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 074468/0229 →
Continuity (4)
Continuation 17700887 · Mar 22, 2022
Provisional Application 63289124 · Dec 13, 2021
Provisional Application 63272648 · Oct 27, 2021
Related Publication 20240414349A1 · Dec 12, 2024
References Cited (23)
US 10142652B2 · Xu · 2018 [cited by examiner]
US 20040264573A1 · Bossen · 2004 [cited by applicant]
US 20130003849A1 · Chien et al. · 2013 [cited by applicant]
US 20180278951A1 · Seregin et al. · 2018 [cited by applicant]
US 20200029091A1 · Chien et al. · 2020 [cited by applicant]
US 20230126552A1 · Zhao · 2023 [cited by examiner]
US 20230156182A1 · Zhao · 2023 [cited by examiner]
US 20230412797A1 · Zhao · 2023 [cited by examiner]
EP 3866470A1 · 2021 [cited by applicant]
JP 2007525100A · 2007 [cited by applicant]
JP 2019533363A · 2019 [cited by applicant]
WO WO2018064524A1 · 2018 [cited by applicant]
WO WO2021040484A1 · 2021 [cited by applicant]
Yue Chen, et al., “An Overview of Core Coding Tools in the AV1 Video Codec”, 978-1-5386-4160-6/18/$31.00 © 2018 IEEE, pp. 41-45. [cited by applicant]
Peter de Rivaz, et al., Argon Design Ltd., “AV1 Bitstream & Decoding Process Specification”, Copyright 2018, The Alliance for Open Media, Last Modified: Jan. 8, 2019 11:48 PT, 681 pages. [cited by applicant]
Elliott Karpilovsky, et al. “Proposal: New Inter Modes for AV2”, Feb. 24, 2021, Alliance for Open Media, Codec Working Group, Document CWG-B018-v1, 6 gages. [cited by applicant]
Benjamin Bross, et al. “Versatile Video Coding Editorial Refinements on Draft 10”, Draft text of video coding specification, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/80 29, Document: JVET-T2… [cited by applicant]
International Search Report and Written Opinion for International Patent Application No. PCT/US22/25054 dated Sep. 1, 2022, 7 pages. [cited by applicant]
Extended European Search Report for European Patent Application No. 22834463.6 dated Oct. 16, 2023, 10 pages. [cited by applicant]
Seregin et al., “Splitting contexts for MVD coding”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JCTVC-J0101, Jul. 12, 2012, 9 pages. [cited by applicant]
Korean-language Office Action issued in Korean Application No. 10-2023-7021020 dated Aug. 14, 2025, with English translation (11 pages). [cited by applicant]
European Office Action issued in European Application No. 22 834 463.6 dated Nov. 27, 2025 (5 pages). [cited by applicant]
Japanese-language Office Action issued in Japanese Application No. 2023-547869 dated Jan. 6, 2026, including English translation (8 pages). [cited by applicant]