IP Library › Granted Patent US 12,739,424
Granted Patent B2
US 12,739,424 · App. 18/932,496 · Granted Sep 15, 2026

Video encoder, video decoder, and corresponding method

Inventors: Huanbang Chen (Shenzhen, CN); Haitao Yang (Shenzhen, CN); Jianle Chen (San Diego, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
H04N19/52H04N19/105H04N19/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,424
App. No.
18/932,496
Granted
Sep 15, 2026
Kind
B2
Abstract

A video encoder, a video decoder, and a corresponding method are provided. The method includes: parsing a bitstream to obtain an index, where the index indicates a target candidate motion vector group of a current coding block; determining the target candidate motion vector group in an affine candidate motion vector list based on the index, where the affine candidate motion vector list includes at least a first candidate motion vector group, the first candidate motion vector group is obtained based on a first group of control points of a first neighboring affine coding block, and the first group of control points is determined based on a CTU located relative to the current coding block, wherein the first neighboring affine coding block is located in the CTU; and predicting a predicted sample value of the current coding block based on the target candidate motion vector group.

Claims (56)

1 . A decoding method, comprising:

in response to an affine inter mode for a current coding block,

parsing a bitstream to obtain an index, wherein the index is used to indicate a target candidate motion vector group of the current coding block;

determining the target candidate motion vector group in a candidate motion vector predictor list based on the index, wherein the target candidate motion vector group represents motion vector predictors of a group of control points of the current coding block, the affine candidate motion vector list comprises at least a first candidate motion vector group, the first candidate motion vector group is obtained based on a first group of control points of a first neighboring affine coding block of the current coding block, and the first group of control points of the first neighboring affine coding block are control points determined based on a relative location of a coding tree unit CTU with respect to the current coding block, wherein the first neighboring affine coding block is located in the CTU;

obtaining a new candidate motion vector group based on motion vector difference MVDs obtained from the bitstream through parsing and the target candidate motion vector group indicated by the index; and

obtaining the motion vectors of the one or more sub-blocks of the current coding block based on the new candidate motion vector group; and

predicting the predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block.

2 . The method according to claim 1 , further comprising:

generating the candidate motion vector predictor MVP list of the current coding block, wherein the candidate motion vector predictor MVP list comprises a plurality of candidate motion vector groups, and the candidate motion vector groups comprise the first candidate motion vector group and a third candidate motion vector group; the first neighboring affine coding block is a lowest coding block in the first CTU, and the current coding block is an uppermost coding block in the second decode tree unit CTU; the third candidate motion vector group is a candidate motion vector predictor of a group of control points of the current coding block obtained by combining candidate motion vectors of at least two control points of the current coding block.

3 . The method according to claim 1 , wherein the parameter model of the current coding block is a 4-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point and a top-right control point of the first neighboring affine coding block.

4 . The method according to claim 1 , wherein the parameter model of the current coding block is a 6-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point, a top-right control point, and a bottom-left control point of the first neighboring affine coding block.

5 . The method according to claim 1 , wherein the predicting a predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block comprises:

predicting the predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block and a reference frame index and a prediction direction that is indicated by the index.

6 . The method according to claim 3 , wherein both location coordinates (x 6 , y 6 ) of a bottom-left control point of the first neighboring affine coding block and location coordinates (x 7 , y 7 ) of a bottom-right control point of the first neighboring affine coding block are derived based on location coordinates (x 4 , y 4 ) of a top-left control point of the first neighboring affine coding block, the location coordinates (x 6 , y 6 ) of the bottom-left control point of the first neighboring affine coding block are represented by (x 4 , y 4 +cuH), and the location coordinates (x 7 , y 7 ) of the bottom-right control point of the first neighboring affine coding block are represented by (x 4 +cuW, y 4 +cuH), wherein cuW is a width of the first neighboring affine coding block, and cuH is a height of the first neighboring affine coding block.

7 . The method according to claim 6 , wherein a motion vector of the bottom-left control point of the first neighboring affine coding block is a motion vector of a bottom-left sub-block of the first neighboring affine coding block, and a motion vector of the bottom-right control point of the first neighboring affine coding block is a motion vector of a bottom-right sub-block of the first neighboring affine coding block.

8 . The method according to claim 1 , wherein the first candidate motion vector group is candidate motion vector predictors of a group of control points of the current coding block obtained based on motion vectors of the first group of control points of the first neighboring affine coding block.

9 . A video data decoding device, comprising:

a memory, configured to store video data in a form of a bitstream; and

a video decoder, configured to:

in response to an affine inter mode for a current coding block, parse the bitstream to obtain an index, wherein the index is used to indicate a target candidate motion vector group of the current coding block;

determine the target candidate motion vector group in a candidate motion vector predictor list based on the index, wherein the target candidate motion vector group represents motion vector predictors of a group of control points of the current coding block, the affine candidate motion vector list comprises at least a first candidate motion vector group, the first candidate motion vector group is obtained based on a first group of control points of a first neighboring affine coding block of the current coding block, and the first group of control points of the first neighboring affine coding block are control points determined based on a relative location of a coding tree unit CTU with respect to the current coding block, wherein the first neighboring affine coding block is located in the CTU;

obtain a new candidate motion vector group based on motion vector difference MVDs obtained from the bitstream through parsing and the target candidate motion vector group indicated by the index; and

obtain the motion vectors of the one or more sub-blocks of the current coding block based on the new candidate motion vector group; and;

predict the predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block.

10 . The video data decoding device according to claim 9 , wherein the video decoder, is further configured to:

generate the candidate motion vector predictor MVP list of the current coding block, wherein the candidate motion vector predictor MVP list comprises a plurality of candidate motion vector groups, and the candidate motion vector groups comprise the first candidate motion vector group and a third candidate motion vector group; the first neighboring affine coding block is a lowest coding block in the first CTU, and the current coding block is an uppermost coding block in the second decode tree unit CTU; the third candidate motion vector group is a candidate motion vector predictor of a group of control points of the current coding block obtained by combining candidate motion vectors of at least two control points of the current coding block.

11 . The video data decoding device according to claim 9 , wherein the parameter model of the current coding block is a 4-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point and a top-right control point of the first neighboring affine coding block.

12 . The video data decoding device according to claim 9 , wherein the parameter model of the current coding block is a 6-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point, a top-right control point, and a bottom-left control point of the first neighboring affine coding block.

13 . The video data decoding device according to claim 9 , wherein the video decoder, is configured to:

predicting the predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block and a reference frame index and a prediction direction that is indicated by the index.

14 . The video data decoding device according to claim 11 , wherein both location coordinates (x 6 , y 6 ) of a bottom-left control point of the first neighboring affine coding block and location coordinates (x 7 , y 7 ) of a bottom-right control point of the first neighboring affine coding block are derived based on location coordinates (x 4 , y 4 ) of a top-left control point of the first neighboring affine coding block, the location coordinates (x 6 , y 6 ) of the bottom-left control point of the first neighboring affine coding block are represented by (x 4 , y 4 +cuH), and the location coordinates (x 7 , y 7 ) of the bottom-right control point of the first neighboring affine coding block are represented by (x 4 +cuW, y 4 +cuH), wherein cuW is a width of the first neighboring affine coding block, and cuH is a height of the first neighboring affine coding block.

15 . The video data decoding device according to claim 14 , wherein a motion vector of the bottom-left control point of the first neighboring affine coding block is a motion vector of a bottom-left sub-block of the first neighboring affine coding block, and a motion vector of the bottom-right control point of the first neighboring affine coding block is a motion vector of a bottom-right sub-block of the first neighboring affine coding block.

16 . The video data decoding device according to claim 9 , wherein the first candidate motion vector group is candidate motion vector predictors of a group of control points of the current coding block obtained based on motion vectors of the first group of control points of the first neighboring affine coding block.

17 . A non-transitory computer-readable media storing computer instructions, that when executed by one or more processors, cause the one or more processors to perform operations, the operations comprising:

in response to an affine inter mode for a current coding block,

parsing a bitstream to obtain an index, wherein the index is used to indicate a target candidate motion vector group of the current coding block;

determining the target candidate motion vector group in a candidate motion vector predictor list based on the index, wherein the target candidate motion vector group represents motion vector predictors of a group of control points of the current coding block, the affine candidate motion vector list comprises at least a first candidate motion vector group, the first candidate motion vector group is obtained based on a first group of control points of a first neighboring affine coding block of the current coding block, and the first group of control points of the first neighboring affine coding block are control points determined based on a relative location of a coding tree unit CTU with respect to the current coding block, wherein the first neighboring affine coding block is located in the CTU;

obtaining a new candidate motion vector group based on motion vector difference MVDs obtained from the bitstream through parsing and the target candidate motion vector group indicated by the index; and

obtaining the motion vectors of the one or more sub-blocks of the current coding block based on the new candidate motion vector group; and;

predicting the predicted sample value of the current coding block based on the motion vectors of the one or more sub-blocks of the current coding block.

18 . The non-transitory computer-readable media according to claim 17 , wherein the operations further comprise:

generating the candidate motion vector predictor MVP list of the current coding block, wherein the candidate motion vector predictor MVP list comprises a plurality of candidate motion vector groups, and the candidate motion vector groups comprise the first candidate motion vector group and a third candidate motion vector group; the first neighboring affine coding block is a lowest coding block in the first CTU, and the current coding block is an uppermost coding block in the second decode tree unit CTU; the third candidate motion vector group is a candidate motion vector predictor of a group of control points of the current coding block obtained by combining candidate motion vectors of at least two control points of the current coding block.

19 . The non-transitory computer-readable media according to claim 17 , wherein the parameter model of the current coding block is a 4-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point and a top-right control point of the first neighboring affine coding block.

20 . The method according to claim 17 , wherein the parameter model of the current coding block is a 6-parameter affine model, and the first group of control points of the first neighboring affine coding block is determined in the following manner:

if the first coding tree unit CTU and the second coding tree unit CTU are in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a bottom-left control point and a bottom-right control point of the first neighboring affine coding block; and;

if the first coding tree unit CTU and the second coding tree unit CTU are not in a top-down relative position relationship, the first group of control points of the first neighboring affine coding block is a top-left control point, a top-right control point, and a bottom-left control point of the first neighboring affine coding block.

Continuity (6)
Continuation 18155641 · Jan 17, 2023
Continuation 17146349 · Jan 11, 2021
Continuation PCTCN2018110436 · Oct 16, 2018
Provisional Application 62737858 · Sep 27, 2018
Provisional Application 62696832 · Jul 11, 2018
Related Publication 20250126287A1 · Apr 17, 2025
References Cited (29)
US 10681370B2 · Chen · 2020 [cited by examiner]
US 10798394B2 · Zhou · 2020 [cited by examiner]
US 11140408B2 · Huang · 2021 [cited by examiner]
US 11223845B2 · Lee · 2022 [cited by examiner]
US 11265573B2 · Liu et al. · 2022 [cited by applicant]
US 11575928B2 · Chen · 2023 [cited by examiner]
US 11653020B2 · Liu et al. · 2023 [cited by applicant]
US 11805272B2 · Poirier · 2023 [cited by examiner]
US 20150296218A1 · Pasupuleti et al. · 2015 [cited by applicant]
CN 104243982A · 2014 [cited by applicant]
CN 104363451A · 2015 [cited by applicant]
CN 104539966A · 2015 [cited by applicant]
CN 104661031A · 2015 [cited by applicant]
CN 104935938A · 2015 [cited by applicant]
CN 105163116A · 2015 [cited by applicant]
EP 3249927A1 · 2017 [cited by applicant]
EP 3331242A1 · 2018 [cited by applicant]
KR 20180019688A · 2018 [cited by applicant]
WO 2014054684A1 · 2014 [cited by applicant]
WO 2017200771A1 · 2017 [cited by applicant]
Study of the affine merge mode; Jul. 2018. (Year: 2018). [cited by examiner]
Vector coding of the Affine MVD; Jul. 2018. (Year: 2018). [cited by examiner]
Minhua Zhou, and Brian Heng, Non-CE4: A study on the affine merge mode, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC1/SC 29/WG 11, JVET-K0052-v2,11th Meeting: Ljubljana,SI,Jul. 7, 2018,pp. 1-10. [cited by applicant]
Seethal Paluri, Mehdi Salehifar, and Seung Hwan Kim, Vector Coding of Affine MVD, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-K0101 (version 3), 11th Meeting:Ljubljana,SI,Jul.… [cited by applicant]
Minhua Zhou, CE4-related: Combined tests of JVET-L0046 and JVET-L0047, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11,JVET-L0048-v1,12th Meeting: Macao,CN,Sep. 2018, pp. 1-11. [cited by applicant]
Huanbang Chen, et al., CE4-related: Combination of affine mode clean up and line buffer reduction,Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-L0694-v3, 12th Meeting:Macao, CN,… [cited by applicant]
Yu Han et al.,“CE4.1.3: Affine motion compensation prediction”, Qualconun Incorporated,Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11,11th Meeting: Ljubljana, SI, Jul. 10-18, 2018,… [cited by applicant]
Z-Y Lin et al:“CE4.1.6:MUP pair list construction for affine inter mode”, 11.JUET Meeting;Jul. 11-Jul. 18, 2018; Ljubljana;(The Joint Uideo Exploration Team of ISO/IEC JTC1/SC29/WG11 and ITU-TSG . 16),No. JVET-K0244,Jul… [cited by applicant]
JVET-K0052-v2, Minhua Zhou et al., Non-CE4: A study on the affine merge mode, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11,11th Meeting: Ljubljana, SI,Jul. 10-18, 2018,10 pages. [cited by applicant]