IP Library Granted Patent US 12,526,420
Granted Patent B2
US 12,526,420 · App. 17/620,402 · Granted Jan 13, 2026

Motion vector prediction in video encoding and decoding

Inventors: Franck Galpin (Thorigne-Fouillard, FR); Fabrice Leleannec (Betton, FR); Antoine Robert (Mézières sur Couesnon, FR); Tangi Poirier (Thorigné-Fouillard, FR)
Assignee: InterDigital CE Patent Holdings, SAS
H04N19/13H04N19/105H04N19/139H04N19/159H04N19/172H04N19/176H04N19/52H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,526,420
App. No.
17/620,402
Granted
Jan 13, 2026
Kind
B2
Abstract

A video codec can involve encoding and decoding picture information and first and second flags, wherein the encoding or decoding of the picture information is based on a coding mode indicated by the first flag or the second flag, and the first flag indicates a subblock merge mode and the second flag indicates an inter affine prediction mode, and the encoding or decoding of the first flag uses Context-Based Adaptive Binary Arithmetic Coding (CABAC) based on a first probability model and encoding or decoding of the second flag uses CABAC based on a second probability model.

Claims (56)

1 . A method comprising:

decoding a first flag included in a bitstream, the first flag being associated with picture information, and the first flag having been encoded using a first Context-Based Arithmetic Coding (CABAC);

decoding a second flag included in the bitstream, the second flag being associated with the picture information, and the second flag having been encoded using a second CABAC separate from the first CABAC; and

decoding at least a portion of the picture information included in the bitstream based on a coding mode indicated by the first flag or the second flag, wherein:

the first flag indicates whether or not a temporal mode is used, wherein in the temporal mode a temporal merge list is constructed that comprises only temporal motion predictor candidates; and

the second flag indicates whether or not an affine prediction mode is used.

2 . The method of claim 1 , wherein:

the affine prediction mode indicates use of an affine motion compensation, and

the temporal merge list comprises a subblock temporal motion vector prediction mode which derives motion information for subblocks of a block from motion information in a reference picture of a picture to which the block belongs.

3 . The method of claim 1 , wherein the temporal merge list comprises at least one of:

a subblock temporal motion vector prediction mode;

a temporal motion vector prediction mode;

a combined inter and intra prediction mode based on a temporal motion vector predictor;

a triangle prediction mode based on a temporal motion vector predictor; or

a merge motion vector difference prediction mode based on a temporal motion vector predictor.

4 . The method of claim 1 , wherein the temporal merge list is placed at one of the following positions in a block decoding process: as a first merge list, after a regular merge list, after an affine merge list, after a combined inter and intra prediction merge list, or after a triangle merge list.

5 . The method of claim 1 , wherein no temporal motion predictor candidate is included in a motion predictor candidate list other than the temporal merge list.

6 . The method of claim 1 , wherein decoding the second flag in a merge mode uses the same second CABAC as decoding the second flag in an inter mode.

7 . The method of claim 1 , wherein the first flag and the second flag are decoded at a block level of the picture information.

8 . A non-transitory computer readable medium storing executable program instructions to cause a computer executing the instructions to perform a method according to claim 1 .

9 . The non-transitory computer readable medium of claim 8 , wherein:

the affine prediction mode indicates use of an affine motion compensation, and

the temporal merge list comprises a subblock temporal motion vector prediction mode which derives motion information for subblocks of a block from motion information in a reference picture of a picture to which the block belongs.

10 . A method comprising:

encoding a first flag associated with picture information, the first flag being encoded using a first Context-Based Arithmetic Coding (CABAC);

encoding a second flag associated with the picture information, the second flag being encoded using a second CABAC separate from the first CABAC; and

encoding at least a portion of the picture information based on a coding mode indicated by the first flag or the second flag to generate a bitstream, wherein:

the first flag indicates whether or not a temporal mode is used, wherein in the temporal mode a temporal merge list is constructed that comprises only temporal motion predictor candidates; and

the second flag indicates whether or not an affine prediction mode is used.

11 . The method of claim 10 , wherein:

the affine prediction mode indicates use of an affine motion compensation, and

the temporal merge list comprises a subblock temporal motion vector prediction mode which derives motion information for subblocks of a block from motion information in a reference picture of a picture to which the block belongs.

12 . The method of claim 10 , wherein the first flag and the second flag are encoded at a block level of the picture information.

13 . An apparatus comprising:

one or more processors configured to:

decode a first flag included in a bitstream, the first flag being associated with picture information, and the first flag having been coded using a first Context-Based Arithmetic Coding (CABAC);

decode a second flag included in the bitstream, the second flag being associated with the picture information, and the second flag having been coded using a second CABAC separate from the first CABAC; and

decode at least a portion of the picture information included in the bitstream based on a coding mode indicated by the first flag or the second flag, wherein:

the first flag indicates whether or not a temporal mode is used, wherein in the temporal mode a temporal merge list is constructed that comprises only temporal motion predictor candidates; and

the second flag indicates whether or not an affine prediction mode is used.

14 . The apparatus of claim 13 , wherein:

the affine prediction mode indicates use of an affine motion compensation, and

the temporal merge list comprises a subblock temporal motion vector prediction mode which derives motion information for subblocks of a block from motion information in a reference picture of a picture to which the block belongs.

15 . The apparatus of claim 13 , wherein decoding the second flag in a merge mode uses the same second CABAC as decoding the second flag in an inter mode.

16 . The apparatus of claim 13 , wherein the first flag and the second flag are decoded at a block level of the picture information.

17 . An apparatus comprising:

one or more processors configured to:

encode a first flag associated with picture information, the first flag being encoded using a first Context-Based Arithmetic Coding (CABAC);

encode a second flag associated with the picture information, the second flag being encoded using a second CABAC separate from the first CABAC; and

encode at least a portion of the picture information based on a coding mode indicated by the first flag or the second flag to generate a bitstream, wherein:

the first flag indicates whether or not a temporal mode is used, wherein in the temporal mode a temporal merge list is constructed that comprises only temporal motion predictor candidates; and

the second flag indicates whether or not an affine prediction mode is used.

18 . The apparatus of claim 17 , wherein:

the affine prediction mode indicates use of an affine motion compensation, and

the temporal merge list comprises a subblock temporal motion vector prediction mode which derives motion information for subblocks of a block from motion information in a reference picture of a picture to which the block belongs.

19 . The apparatus of claim 17 , wherein the first flag and the second flag are encoded at a block level of the picture information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: INTERDIGITAL VC HOLDINGS FRANCE, SAS
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 064460/0921 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2021
From: GALPIN, FRANCK; LELEANNEC, FABRICE; ROBERT, ANTOINE; POIRIER, TANGI
To: INTERDIGITAL VC HOLDINGS FRANCE, SAS
Reel/Frame 058419/0102 →
Priority Claims (1)
EP 19305833 · Jun 25, 2019 · regional
Continuity (1)
Related Publication 20230018401A1 · Jan 19, 2023
References Cited (37)
US 9930363B2 · Rusanovskyy et al. · 2018 [cited by applicant]
US 10148977B2 · Yu et al. · 2018 [cited by applicant]
US 11438578B2 · Chen et al. · 2022 [cited by applicant]
US 12069239B2 · Zhang · 2024 [cited by examiner]
US 20130003827A1 · Misra et al. · 2013 [cited by applicant]
US 20140029670A1 · Kung et al. · 2014 [cited by applicant]
US 20170332095A1 · Zou et al. · 2017 [cited by applicant]
US 20200288141A1 · Chono · 2020 [cited by applicant]
US 20210160533A1 · Zhang · 2021 [cited by examiner]
US 20210392329A1 · Galpin · 2021 [cited by examiner]
US 20220053204A1 · Laroche · 2022 [cited by examiner]
US 20220070444A1 · Takehara · 2022 [cited by examiner]
US 20220086433A1 · Zhang · 2022 [cited by examiner]
US 20230396775A1 · Park · 2023 [cited by examiner]
CN 105308965A · 2016 [cited by applicant]
CN 107534711A · 2018 [cited by applicant]
CN 109547790A · 2019 [cited by applicant]
JP 2022511637A · 2022 [cited by applicant]
RU 2645270C1 · 2018 [cited by applicant]
WO 2014120721A1 · 2014 [cited by applicant]
WO 2019069602A1 · 2019 [cited by applicant]
WO 2020085953A1 · 2020 [cited by applicant]
WO WO2020077003A1 · 2020 [cited by applicant]
Galpin et al., “CE4-related: simplified constructed temporal affine merge candidates”, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC 1/SC 29/WG11, Document: JVET-L0522-v2, 12th Meeting: Macao, China,… [cited by applicant]
Chen et al., “Context Reduction for Inter and Split Syntax Elements”, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC 1/SC29/WG11, Document: JVET-N0600, 14th Meeting: Geneva, Switzerland, Mar. 19, 2019… [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 7 (JEM 7)”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document JVET-G1001-v1, 7th Meeting: Torino, Italy, … [cited by applicant]
Galpin et al., “Non-CE4: Affine and sub-block modes coding clean-up”, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC 1/SC29/WG11, Document: JVET-O0500, 15th Meeting: Gothenburg, Sweden, Jul. 3, 2019, … [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video—Reference software for ITU-T H.265 high efficiency video coding”, International Telecommunication U… [cited by applicant]
Chen et al, “Algorithm description for Versatile Video Coding and Test Model 5 (VTM 5)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IECJTC 1/SC 29/WG 11, Document: JVET-N1002-v2, 14th Meeting, Geneva, S… [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 5) ”, Joint Video Experts Team (JVET) of ISO/IEC JTC1/SC29/WG11 and ITU-T SG.16), Document: JVET-N1001-v8, 14th Meeting: Geneva, Switzerland, Mar. 19, 2019, 400 pages. [cited by applicant]
Bross, Benjamin et al., “Versatile Video Coding (Draft 3)”, JVET-L1001-v9, Editors, JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, 233 pages. [cited by applicant]
Marpe, et al., “Context-Based Adaptive Binary Arithmetic Coding in the H.264/AVC Video Compression Standard”, IEEE Transactions on Circuits and Systems for Video Technology, vol. X, No. Y, May 21, 2003, 18 pages. [cited by applicant]
Lai et al., “CE8-related: Clarification on Interaction Between CPR and Other Inter Coding Tools”, JVET-M0175-v1, MediaTek Inc., Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Mee… [cited by applicant]
Laroche et al., “CE4-related: On Merge Index Coding”, JVET-L0194, Canon, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, pp. 1-3. [cited by applicant]
Li et al., “CE4-related: Affine Merge Mode with Prediction Offsets”, JVET-L0320, Tencent America, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, China, Oct. 3-12,… [cited by applicant]
Yang et al., “CE4: Summary Report on Inter Prediction and Motion Vector Coding”, JVET-L0024-v1, CE4 Coordinators, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, C… [cited by applicant]
Xiu, et al., “Draft text for advanced temporal motion vector prediction (ATMVP) and adaptive motion vector resolution (AMVR)”, JVET-K0566_r3 Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG… [cited by applicant]