IP Library Granted Patent US 12,464,160
Granted Patent B2
US 12,464,160 · App. 16/753,763 · Granted Nov 4, 2025

Methods and apparatuses for video encoding and video decoding

Inventors: Antoine Robert (Cesson-Sevigne, FR); Fabrice Leleannec (Cesson-Sevigne, FR); Tangi Poirier (Cesson-Sevigne, FR)
Assignee: InterDigital VC Holdings, Inc.
H04N19/56H04N19/105H04N19/139H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,464,160
App. No.
16/753,763
Granted
Nov 4, 2025
Kind
B2
Abstract

Implementations are described for determining, for a block being encoded in a picture, at least one predictor candidate, determining for the at least one predictor candidate, one or more corresponding control point generator motion vectors, based on motion information associated to the at least one predictor candidate, determining for the block being encoded, one or more corresponding control point motion vectors, based on the one or more corresponding control point generator motion vectors determined for the at least one predictor candidate, determining, based on the one or more corresponding control point motion vectors determined for the block, a corresponding motion field, and encoding the block based on the corresponding motion field.

Claims (63)

1 . A method for video decoding, comprising:

determining a predictor candidate for a block being decoded in a picture in an affine motion model, wherein the predictor candidate has a translational motion model and a plurality of sub-blocks comprising at least a top-left sub-block, a top-right sub-block, a bottom-left sub-block, and a bottom-right sub-block;

determining for the predictor candidate, at least two control point generator motion vectors of an affine motion model, wherein each control point generator motion vector is associated to a different sub-block of the predictor candidate, provided that the at least two control point generator motion vectors determined for the predictor candidate are for the top-left sub-block and the top-right sub-block, respectively, and wherein motion vectors for the bottom-left sub-block and the bottom-right sub-block are compared to estimated motion vectors for the bottom-left sub-block and the bottom-right sub-block and satisfy a threshold level for respective angle and magnitude;

determining corresponding control point motion vectors for the block being decoded based on the at least two control point generator motion vectors determined for the predictor candidate, such that the determined control point motion vectors reflect both motion per sub-block and the translational motion model of the predictor candidate;

determining, based on the determined control point motion vectors, a corresponding motion field for the block, wherein the motion field identifies motion vectors used for prediction of sub-blocks of the block being decoded; and

decoding the block based on the motion field.

2 . The method of claim 1 , wherein the predictor candidate is comprised in a set of predictor candidates and wherein determining the predictor candidate comprises receiving an index corresponding to the predictor candidate in the set of predictor candidates.

3 . The method of claim 1 , further comprising verifying that the determined control point motion vectors satisfy the affine motion model.

4 . The method of claim 1 , wherein determining the at least two control point generator motion vectors comprises:

determining, for at least two distinct sets of at least three sub-blocks of the predictor candidate, corresponding control point motion vectors for the predictor candidate associated respectively to the at least two sets, based on the motion vectors associated respectively to the at least three sub-blocks of each set; and

calculating corresponding control point motion vectors associated to the predictor candidate by averaging the determined control point motion vectors associated to each set.

5 . The method of claim 1 , wherein the motion vector is derived from at least one of:

a bilateral template matching between two reference blocks in respectively two reference frames;

a reference block of a reference frame identified by motion information of a first spatial neighboring block of the predictor candidate; or

an average of motion vectors of spatial and temporal neighboring blocks of the predictor candidate.

6 . A non-transitory computer readable storage medium having stored thereon instructions for decoding video data according to the method of claim 1 .

7 . An apparatus for video decoding, comprising a memory and at least one processor configured for:

determining a predictor candidate for a block being decoded in a picture in an affine motion model, wherein the predictor candidate has a translational motion model and a plurality of sub-blocks comprising at least a top-left sub-block, a top-right sub-block, a bottom-left sub-block, and a bottom-right sub-block;

determining for the predictor candidate, at least two control point generator motion vectors of an affine motion model, wherein each control point generator motion vector is associated to a different sub-block of the predictor candidate, provided that the at least two control point generator motion vectors determined for the predictor candidate are for the top-left sub-block and the top-right sub-block, respectively, and wherein motion vectors for the bottom-left sub-block and the bottom-right sub-block are compared to estimated motion vectors for the bottom-left sub-block and the bottom-right sub-block and satisfy a threshold level for respective angle and magnitude;

determining corresponding control point motion vectors for the block being decoded based on the at least two control point generator motion vectors determined for the predictor candidate, such that the determined control point motion vectors reflect both motion per sub-block and the translational motion model of the predictor candidate;

determining, based on the determined control point motion vectors, a corresponding motion field for the block, wherein the motion field identifies motion vectors used for prediction of sub-blocks of the block being decoded; and

decoding the block based on the motion field.

8 . The apparatus of claim 7 , wherein the predictor candidate is comprised in a set of predictor candidates and wherein determining the predictor candidate comprises receiving an index corresponding to the predictor candidate in the set of predictor candidates.

9 . The apparatus of claim 7 , wherein the at least one processor is further configured for verifying that the determined control point motion vectors satisfy the affine motion model.

10 . The apparatus of claim 7 , wherein determining the at least two control point generator motion vectors comprises:

determining, for at least two distinct sets of at least three sub-blocks of the predictor candidate, corresponding control point motion vectors for the predictor candidate associated respectively to the at least two sets, based on the motion vectors associated respectively to the at least three sub-blocks of each set; and

calculating corresponding control point motion vectors associated to the predictor candidate by averaging the determined control point motion vectors associated to each set.

11 . The apparatus of claim 7 , wherein the motion vector is derived from at least one of:

a bilateral template matching between two reference blocks in respectively two reference frames;

a reference block of a reference frame identified by motion information of a first spatial neighboring block of the predictor candidate; or

an average of motion vectors of spatial and temporal neighboring blocks of the predictor candidate.

12 . A method for video encoding, comprising:

determining a predictor candidate for a block being encoded in a picture in an affine motion model, wherein the predictor candidate has a translational motion model and a plurality of sub-blocks comprising at least a top-left sub-block, a top-right sub-block, a bottom-left sub-block, and a bottom-right sub-block;

determining for the predictor candidate, at least two control point generator motion vectors of an affine motion model, wherein each control point generator motion vector is associated to a different sub-block of the predictor candidate, provided that the at least two control point generator motion vectors determined for the predictor candidate are for the top-left sub-block and the top-right sub-block, respectively, and wherein motion vectors for the bottom-left sub-block and the bottom-right sub-block are compared to estimated motion vectors for the bottom-left sub-block and the bottom-right sub-block and satisfy a threshold level for respective angle and magnitude;

determining corresponding control point motion vectors for the block being encoded based on the at least two control point generator motion vectors determined for the predictor candidate, such that the determined control point motion vectors reflect both motion per sub-block and the translational motion model of the predictor candidate;

determining, based on the determined control point motion vectors, a corresponding motion field for the block, wherein the motion field identifies motion vectors used for prediction of sub-blocks of the block being encoded; and

encoding the block based on the motion field.

13 . The encoding method of claim 12 , wherein the predictor candidate is comprised in a set of predictor candidates, the encoding method further comprising:

encoding an index for the predictor candidate from the set of predictor candidates.

14 . The method of claim 12 , further comprising verifying that the determined control point motion vectors satisfy the affine motion model.

15 . The method of claim 12 , wherein determining the at least two control point generator motion vectors comprises:

determining, for at least two distinct sets of at least three sub-blocks of the predictor candidate, corresponding control point motion vectors for the predictor candidate associated respectively to the at least two sets, based on the motion vectors associated respectively to the at least three sub-blocks of each set; and

calculating corresponding control point motion vectors associated to the predictor candidate by averaging the determined control point motion vectors associated to each set.

16 . The method of claim 12 , wherein the motion vector is derived from at least one of:

a bilateral template matching between two reference blocks in respectively two reference frames;

a reference block of a reference frame identified by motion information of a first spatial neighboring block of the predictor candidate; or

an average of motion vectors of spatial and temporal neighboring blocks of the predictor candidate.

17 . A non-transitory computer readable storage medium having stored thereon instructions for encoding video data according to the method of claim 12 .

18 . An apparatus for video encoding, comprising a memory and at least one processor configured for:

determining a predictor candidate for a block being encoded in a picture in an affine motion model, wherein the predictor candidate has a translational motion model and a plurality of sub-blocks comprising at least a top-left sub-block, a top-right sub-block, a bottom-left sub-block, and a bottom-right sub-block;

determining for the predictor candidate, at least two control point generator motion vectors of an affine motion model, wherein each control point generator motion vector is associated to a different sub-block of the predictor candidate, provided that the at least two control point generator motion vectors determined for the predictor candidate are for the top-left sub-block and the top-right sub-block respectively, and wherein motion vectors for the bottom-left sub-block and the bottom-right sub-block are compared to estimated motion vectors for the bottom-left sub-block and the bottom-right sub-block and satisfy a threshold level for respective angle and magnitude;

determining corresponding control point motion vectors for the block being encoded based on the at least two control point generator motion vectors determined for the predictor candidate, such that the determined control point motion vectors reflect both motion per sub-block and the translational motion model of the predictor candidate;

determining, based on the determined control point motion vectors, a corresponding motion field for the block, wherein the motion field identifies motion vectors used for prediction of sub-blocks of the block being encoded; and

encoding the block based on the motion field.

19 . The apparatus of claim 18 , wherein the predictor candidate is comprised in a set of predictor candidates, and the at least one processor is further configured for encoding an index for the predictor candidate from the set of predictor candidates.

20 . The apparatus of claim 18 , wherein the at least one processor is further configured for verifying that the one or more corresponding control point motion vectors associated to the predictor candidate satisfies the affine motion model.

21 . The apparatus of claim 18 , wherein determining the at least two control point generator motion vectors comprises:

determining, for at least two distinct sets of at least three sub-blocks of the predictor candidate, corresponding control point motion vectors for the predictor candidate associated respectively to the at least two sets, based on the motion vectors associated respectively to the at least three sub-blocks of each set; and

calculating corresponding control point motion vectors associated to the predictor candidate by averaging the determined control point motion vectors associated to each set.

22 . The apparatus of claim 18 , wherein the motion vector is derived from at least one of:

a bilateral template matching between two reference blocks in respectively two reference frames;

a reference block of a reference frame identified by motion information of a first spatial neighboring block of the predictor candidate; or

an average of motion vectors of spatial and temporal neighboring blocks of the predictor candidate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2020
From: ROBERT, ANTOINE; LELEANNEC, FABRICE; POIRIER, TANGI
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 052666/0184 →
Priority Claims (1)
EP 17306336 · Oct 5, 2017 · regional
Continuity (1)
Related Publication 20200288166A1 · Sep 10, 2020
References Cited (26)
US 20140241434A1 · Lin et al. · 2014 [cited by applicant]
US 20160337662A1 · Pang et al. · 2016 [cited by applicant]
US 20170332095A1 · Zou · 2017 [cited by examiner]
US 20180014017A1 · Li · 2018 [cited by examiner]
US 20180324454A1 · Lin · 2018 [cited by examiner]
US 20190082191A1 · Chuang · 2019 [cited by examiner]
CN 103907346A · 2014 [cited by applicant]
CN 104935938A · 2015 [cited by applicant]
JP 2004364333A · 2004 [cited by applicant]
TW 201701671A · 2017 [cited by applicant]
WO 2017118409A1 · 2017 [cited by applicant]
WO WO2017148345A1 · 2017 [cited by applicant]
WO WO2017157259A1 · 2017 [cited by applicant]
Karczewicz et al., “JVET AHG report: Tool evaluation (AHG1)”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document JVET-G0001, 7th meeting, Torino, Italy, Jul. 13, 2017, 6 page… [cited by applicant]
Huang et al., “Affine SKIP and DIRECT Modes for Efficient Video Coding”, 2012 Conference on Visual Communications and Image Processing, San Diego, California, USA, Nov. 27, 2012, 6 pages. [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 2”, Joint Video Exploration Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Document: JVET-B1001 v3, 2nd Meeting, San Diego, California, USA,… [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 6 (JEM 6)”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document JVET-F1001-v2, 6th meeting, Hobart, Austral… [cited by applicant]
Anonymous, “Reference software for ITU-T H.265 high efficiency video coding”, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, I… [cited by applicant]
Anonymous, “Affine transform prediction for next generation video coding”, Study Group 16—Contribution 1016, Huawei Technologies Co., Ltd., International Telecommunication Union Telecommunication Standardization Sector,… [cited by applicant]
Li et al., “An Affine Motion Compensation Framework for High Efficiency Video Coding”, 2015 IEEE International Symposium on Circuits and Systems (ISCAS), Lisbon, Portugal, May 24, 2015, pp. 525-528. [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 5 (JEM 5)”, JVET-E1001-V2, Editors, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 5th Meeting: Geneva, CH, Jan… [cited by applicant]
English Language Translation, Chinese Publication No. CN 104935938 A. [cited by applicant]
English Language Translation, Japanese Publication No. JP 2004364333 A. [cited by applicant]
Chen, et al., “Algorithm Description of Joint Exploration Test Model 6 (JEM 6)”, JVET-F1001-V3, Editors, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 6th Meeting: Hobart, AU, Mar… [cited by applicant]
Huawei Technologies, “Affine transform prediction for next generation video coding”, ITU-T SG16 Meeting; Dec. 10, 2015-Oct. 23, 2015; Geneva, No. T13-SG16-C-1016, XP030100743, Sep. 29, 2015, pp. 1-11. [cited by applicant]
Chen, et al., “Description of SDR, HDR and 360° video coding technology proposal by Qualcomm and Technicolor—low and high complexity versions”, JVET of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11; 10th Meeting: San D… [cited by applicant]
Cited By (1)
US 12,666,047