IP Library Granted Patent US 12,401,795
Granted Patent B2
US 12,401,795 · App. 17/832,871 · Granted Aug 26, 2025

Video data inter prediction based on control points

Inventors: Huanbang Chen (Shenzhen, CN); Shan Gao (Dongguan, CN); Haitao Yang (Shenzhen, CN)
Assignee: Huawei Technologies Co., Ltd.
H04N19/137H04N19/105H04N19/132H04N19/159H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,401,795
App. No.
17/832,871
Granted
Aug 26, 2025
Kind
B2
Abstract

A video data inter prediction method is provided, which includes: determining a candidate motion information list of a current picture block, where the candidate motion information list includes at least one first candidate motion information group, at least one second candidate motion information group, the first candidate motion information group is a motion information group determined based on motion information of preset locations on a first neighboring picture block of the current picture block and a motion model of the first neighboring picture block, the second candidate motion information group is a set of motion information of at least two sample locations that are respectively neighboring to at least two preset locations on the current picture block; determining target motion information from the candidate motion information list; and performing inter prediction on the current picture block based on the target motion information.

Claims (59)

1. A method, comprising:

determining a candidate motion information list of a current picture block to be coded using an affine merge mode, the candidate motion information list comprising at least one first candidate motion information group and at least one second candidate motion information group both comprising motion vectors for a first group of at least two control points of the current picture block, wherein a plurality of pieces of index information indicate the first candidate motion information group and the second candidate motion information group respectively, the first candidate motion information group is a motion information group determined based on motion information of control points of a first affine neighboring picture block of the current picture block and a motion model of the first affine neighboring picture block, and the second candidate motion information group is based on a set of motion information of at least two sample locations that are respectively neighboring to control points of a second group of at least two control points of the current picture block, and the at least two sample locations are located on at least one second neighboring picture block of the current picture block;

determining target motion information from the candidate motion information list; and

performing inter prediction on the current picture block based on the target motion information; and

wherein the determining the candidate motion information list of the current picture block comprises:

in response to that the second group of at least two control points is different from the first group of at least two control points, deriving the second candidate motion information group according to a location transformation formula and based on motion information corresponding to the second group of at least two control points of the current picture block.

2. The method according to claim 1 , wherein determining the candidate motion information list of the current picture block comprises:

in response to that the first affine neighboring picture block is a picture block using an affine motion model, deriving, based on motion information of at least two control points of the first affine neighboring picture block and the affine motion model of the first affine neighboring picture block, motion information of at least two control points of the current picture block, and adding the motion information of the at least two control points of the current picture block into the candidate motion information list as the first candidate motion information group, wherein the at least two control points of the first affine neighboring picture block corresponds to the at least two control points of the current picture block.

3. The method according to claim 1 , wherein an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are four-parameter affine motion models; or

an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are six-parameter affine motion models.

4. The method according to claim 1 , wherein determining the candidate motion information list of the current picture block comprises:

adding the at least one first candidate motion information group to the candidate motion information list, and subsequently adding the at least one second candidate motion information group to the candidate motion information list, and

adding zero motion information to the candidate motion information list in response to that a length of the candidate motion information list is less than a predefined length N upon having added the at least one first candidate motion information group and the at least one second candidate motion information group.

5. The method according to claim 1 , wherein the location transformation formula is based on a 4-parameter affine model or a 6-parameter model and coordinates of the first group of control points and the second group of control points.

6. The method according to claim 1 , wherein:

the second group of at least two control points is a top-left control point, a top-right control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, the top-right control point, and a bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, a top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-right control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is a top-left control point, the top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point and a bottom-left control point of the current picture block and the first group of at least two control points is the top-left control point and a top-right control point of the current picture block.

7. A device, comprising:

one or more processors; and

a non-transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, wherein the programming, when executed by the one or more processors, causes the device to:

determine a candidate motion information list of a current picture block to be coded using an affine merge mode, wherein the candidate motion information list comprises at least one first candidate motion information group and at least one second candidate motion information group both comprising motion vectors for a first group of at least two control points of the current picture block, wherein a plurality of pieces of index information indicate the first candidate motion information group and the second candidate motion information group respectively, the first candidate motion information group is a motion information group determined based on motion information of control points of a first affine neighboring picture block of the current picture block and a motion model of the first affine neighboring picture block, and the second candidate motion information group is based on a set of motion information of at least two sample locations that are respectively neighboring to control points of a second group of at least two control points of the current picture block, and the at least two sample locations are located on at least one second neighboring picture block of the current picture block;

determine target motion information from the candidate motion information list; and

perform inter prediction on the current picture block based on the target motion information; and

wherein the determining the candidate motion information list of the current picture block comprises:

in response to that the second group of at least two control points is different from the first group of at least two control points, deriving the second candidate motion information group according to a location transformation formula and based on motion information corresponding to the second group of at least two control points.

8. The device according to claim 7 , wherein the programming, when executed by the one or more processors, causes the device to determine the candidate motion information list of the current picture block by:

in response to that the first affine neighboring picture block is a picture block using an affine motion model, deriving, based on motion information of at least two control points of the first affine neighboring picture block and the affine motion model of the first affine neighboring picture block, motion information of at least two control points of the current picture block, and adding the motion information of the at least two control points of the current picture block into the candidate motion information list as the first candidate motion information group, wherein the at least two control points of the first affine neighboring picture block corresponds to the at least two control points of the current picture block.

9. The device according to claim 7 , wherein an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are four-parameter affine motion models; or

an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are six-parameter affine motion models.

10. The device according to claim 7 , wherein the programming, when executed by the one or more processors, causes the device to determine the candidate motion information list of the current picture block by:

adding the at least one first candidate motion information group to the candidate motion information list, and subsequently adding the at least one second candidate motion information group to the candidate motion information list; and

adding zero motion information to the candidate motion information list in response to that a length of the candidate motion information list is less than a predefined length N upon having added the at least one first candidate motion information group and the at least one second candidate motion information group.

11. The device according to claim 7 , wherein the location transformation formula is based on a 4-parameter affine model or a 6-parameter model and coordinates of the first group of control points and the second group of control points.

12. The device according to claim 7 , wherein:

the second group of at least two control points is a top-left control point, a top-right control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, the top-right control point, and a bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, a top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-right control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is a top-left control point, the top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point and a bottom-left control point of the current picture block and the first group of at least two control points is the top-left control point and a top-right control point of the current picture block.

13. A non-transitory computer-readable medium carrying a program code which, when executed by a computer device, causes the computer device to perform steps comprising:

determining a candidate motion information list of a current picture block to be coded using an affine merge mode, the candidate motion information list comprising at least one first candidate motion information group and at least one second candidate motion information group both comprising motion vectors for a first group of at least two control points of the current picture block, wherein a plurality of pieces of index information indicate the first candidate motion information group and the second candidate motion information group respectively, the first candidate motion information group is a motion information group determined based on motion information of control points of a first affine neighboring picture block of the current picture block and a motion model of the first affine neighboring picture block, and the second candidate motion information group is based on a set of motion information of at least two sample locations that are respectively neighboring to control points of a second group of at least two control points of the current picture block, and the at least two sample locations are located on at least one second neighboring picture block of the current picture block;

determining target motion information from the candidate motion information list; and

performing inter prediction on the current picture block based on the target motion information; and

wherein the determining the candidate motion information list of the current picture block comprises:

in response to that the second group of at least two control points is different from the first group of at least two control points, deriving the second candidate motion information group according to a location transformation formula and based on motion information corresponding to the second group of at least two control points.

14. The non-transitory computer-readable medium according to claim 13 , wherein the determining the candidate motion information list of the current picture block comprises:

when the first affine neighboring picture block is a picture block using an affine motion model, deriving, based on motion information of at least two control points of the first affine neighboring picture block and the affine motion model of the first affine neighboring picture block, motion information of at least two control points of the current picture block, and adding the motion information of the at least two control points of the current picture block into the candidate motion information list as the first candidate motion information group, wherein the at least two control points of the first affine neighboring picture block corresponds to the at least two control points of the current picture block.

15. The non-transitory computer-readable medium according to claim 13 , wherein an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are four-parameter affine motion models; or

an affine motion model of the current picture block and an affine motion model of the first affine neighboring picture block are six-parameter affine motion models.

16. The non-transitory computer-readable medium according to claim 13 , wherein determining the candidate motion information list of the current picture block comprises:

adding the at least one first candidate motion information group to the candidate motion information list, and subsequently adding the at least one second candidate motion information group to the candidate motion information list, and

adding zero motion information to the candidate motion information list in response to that a length of the candidate motion information list is less than a predefined length N upon having added the at least one first candidate motion information group and the at least one second candidate motion information group.

17. The non-transitory computer-readable medium according to claim 13 , wherein the location transformation formula is based on a 4-parameter affine model or a 6-parameter model and coordinates of the first group of control points and the second group of control points.

18. The non-transitory computer-readable medium according to claim 13 , wherein:

the second group of at least two control points is a top-left control point, a top-right control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, the top-right control point, and a bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is the top-left control point, a top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-right control point, a bottom-left control point, and a bottom-right control point of the current picture block and the first group of at least two control points is a top-left control point, the top-right control point, and the bottom-left control point of the current picture block; or

the second group of at least two control points is a top-left control point and a bottom-left control point of the current picture block and the first group of at least two control points is the top-left control point and a top-right control point of the current picture block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2024
From: CHEN, HUANBANG; GAO, SHAN; YANG, HAITAO
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 067748/0732 →
Priority Claims (1)
CN 201711319298.2 · Dec 12, 2017 · national
Continuity (3)
Continuation 16898973 · Jun 11, 2020
Continuation PCTCN2018120435 · Dec 12, 2018
Related Publication 20220394270A1 · Dec 8, 2022
References Cited (42)
US 10462488B1 · Li · 2019 [cited by examiner]
US 10602180B2 · Chen et al. · 2020 [cited by applicant]
US 10701390B2 · Li · 2020 [cited by examiner]
US 10728571B2 · Son et al. · 2020 [cited by applicant]
US 10798394B2 · Zhou · 2020 [cited by applicant]
US 20170332095A1 · Zou · 2017 [cited by examiner]
US 20180098063A1 · Chen · 2018 [cited by examiner]
US 20190028731A1 · Chuang · 2019 [cited by examiner]
US 20190208211A1 · Zhang · 2019 [cited by examiner]
US 20190222834A1 · Chen et al. · 2019 [cited by applicant]
US 20190222865A1 · Zhang et al. · 2019 [cited by applicant]
US 20190364295A1 · Li · 2019 [cited by examiner]
US 20200021836A1 · Xu et al. · 2020 [cited by applicant]
US 20200059651A1 · Lin · 2020 [cited by examiner]
US 20200288163A1 · Poirier et al. · 2020 [cited by applicant]
US 20200288166A1 · Robert et al. · 2020 [cited by applicant]
US 20210051339A1 · Liu et al. · 2021 [cited by applicant]
US 20220394270A1 · Chen · 2022 [cited by examiner]
AU 2017340631A1 · 2019 [cited by applicant]
CN 1325220A · 2001 [cited by applicant]
CN 103716631A · 2014 [cited by applicant]
CN 104539966A · 2015 [cited by applicant]
CN 104935938A · 2015 [cited by applicant]
CN 107046645A · 2017 [cited by applicant]
CN 107071462A · 2017 [cited by applicant]
JP 2012165279A · 2012 [cited by applicant]
KR 100331048B1 · 2002 [cited by applicant]
TW I486061B · 2015 [cited by applicant]
WO 2017022973A1 · 2017 [cited by applicant]
WO 2017026681A1 · 2017 [cited by applicant]
WO 2017118409A1 · 2017 [cited by applicant]
WO 2017118411A1 · 2017 [cited by applicant]
WO 2017147765A1 · 2017 [cited by applicant]
WO 2017148345A1 · 2017 [cited by applicant]
Chen et al., “Description of SDR, HDR and 360° video coding technology proposal by Huawei, GoPro, HiSilicon, and Samsung,” buJoint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-J0… [cited by applicant]
Office Action in Japanese Appln. No. 2022-151544, mailed on Dec. 5, 2023, 6 pages (with English translation). [cited by applicant]
Jianle Chen et al. Algorithm Description of Joint Exploration Test Model 7 (JEM 7), Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-G1001-v1, 7th Meeting: Torino, IT, Jul. 13-… [cited by applicant]
Cordula Heithausen et al,“Inter Prediction using Estimation and Explicit Coding of Affine Parameters”,Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 8th Meeting: Macao, CN, Oct. 1… [cited by applicant]
Marta Karczewicz et al,“JVET AHG report: Tool evaluation (AHG1)”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, Document: JVET-G0001, t… [cited by applicant]
Kwak, “A Prediction Search Algorithm in Video Coding by using Neighboring-Block Motion Vectors,” Journal of the Korea Academia-Industrial cooperation Society, 2011, vol. 12, No. 8, pp. 3697-3705 (with English abstract). [cited by applicant]
Office Action in Australian Appln. No. 2023200956, mailed on Apr. 19, 2024, 3 pages. [cited by applicant]
Akula et al., Joint Video Exploration Team (JVET), “Description of SDR, HDR and 360° video coding technology proposal considering mobile application scenario by Samsung, Huawei, GoPro, and HiSilicon,” Samsung Electronic… [cited by applicant]