IP Library › Granted Patent US 12,425,643
Granted Patent B2
US 12,425,643 · App. 18/188,908 · Granted Sep 23, 2025

Derivation of affine merge candidates with linear regression for video coding

Inventors: Yan Zhang (San Diego, CA); Han Huang (San Diego, CA); Vadim Seregin (San Diego, CA); Muhammed Zeyd Coban (Carlsbad, CA); Marta Karczewicz (San Diego, CA)
Assignee: QUALCOMM INCORPORATED
H04N19/573H04N19/105H04N19/139H04N19/159H04N19/176H04N19/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,425,643
App. No.
18/188,908
Granted
Sep 23, 2025
Kind
B2
Abstract

A method for coding a block of video data using affine mode includes determining a refined affine model for the current block of video data from a linear regression process using a base motion vector field and a guidance motion vector field as inputs to the linear regression process. The method further includes determining affine merge candidates for the current block using the refined affine model, and coding the current block of video data using the affine merge candidates.

Claims (128)

1. A method of decoding video data, the method comprising:

receiving a current block of video data to be decoded using an affine merge mode;

determining a base motion vector field from an affine model associated with a neighboring block;

determining a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

determining a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

determining affine merge candidates for the current block using the refined affine model; and

decoding the current block of video data using the affine merge candidates.

2. The method of claim 1 , wherein the plurality of sub-blocks adjacent to the current block include a first plurality of sub-blocks in a row adjacent to the current block and a second plurality of sub-blocks in a column adjacent to the current block.

3. The method of claim 1 , wherein the neighboring block is adjacent to the current block of video data.

4. The method of claim 1 , wherein the neighboring block is not adjacent to the current block of video data.

5. The method of claim 4 , wherein the affine model associated with the neighboring block is stored in a history table.

6. The method of claim 5 , further comprising:

determining that the neighboring block is at a position that is outside of a distance constraint of the history table;

converting a coordinate of the neighboring block to be an adjusted coordinate located within the distance constraint of the history table; and

determining the affine model from a second affine model associated with a second block located at the adjusted coordinate.

7. The method of claim 6 , further comprising:

determining the base motion vector field from a subset of sub-blocks of the second block based on the second block not being fully within the distance constraint of the history table.

8. The method of claim 1 , further comprising:

determining the base motion vector field from a neighboring block coded using affine mode; and

determining the guidance motion vector field from motion information associated with a subset of a plurality of sub-blocks adjacent to the current block.

9. The method of claim 8 , further comprising:

determining the subset of the plurality of sub-blocks adjacent to the current block based on a location of the neighboring block coded using the affine mode.

10. The method of claim 8 , wherein the subset of the plurality of sub-blocks adjacent to the current block is one of a row of sub-blocks adjacent to the current block or a column of sub-blocks adjacent to the current block.

11. The method of claim 8 , wherein the subset of the plurality of sub-blocks adjacent to the current block are within a distance constraint from the current block.

12. The method of claim 8 , wherein determining the guidance motion vector field from motion information associated with the subset of the plurality of sub-blocks adjacent to the current block comprises:

sub-sampling the plurality of sub-blocks adjacent to the current block.

13. The method of claim 1 , wherein the guidance motion vector field is a sub-block temporal motion vector predictor field.

14. The method of claim 1 , further comprising:

determining the base motion vector field by performing a linear regression process on two motion vector fields.

15. The method of claim 1 , wherein the base motion vector field includes motion information from sub-blocks that are adjacent to the current block and motion information from sub-blocks that are not adjacent to the current block.

16. The method of claim 1 , wherein determining the refined affine model for the current block of video data comprises:

determining the refined affine model using the base motion vector field, the guidance motion vector field, and at least one other motion vector field as inputs to a linear regression process.

17. The method of claim 1 , further comprising:

scaling motion vectors in one or more of the base motion vector field or the guidance motion vector field relative to a most frequently used reference picture index.

18. The method of claim 1 , further comprising:

converting a uni-directional affine merge candidate of the affine merge candidates to a bi-directional affine merge candidate.

19. The method of claim 1 , the method further comprising:

receiving an index that specifies a particular affine merge candidate of the affine merge candidates;

determining a motion vector for the current block based on the particular affine merge candidate; and

decoding the current block using the motion vector.

20. An apparatus configured to decode video data, the apparatus comprising:

a memory configured to store video data; and

one or more processors implemented in circuitry and in communication with the memory, the one or more processors configured to:

receive a current block of video data to be decoded using an affine merge mode;

determine a base motion vector field from an affine model associated with a neighboring block;

determine a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

determine a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

determine affine merge candidates for the current block using the refined affine model; and

decode the current block of video data using the affine merge candidates.

21. The apparatus of claim 20 , wherein the plurality of sub-blocks adjacent to the current block include a first plurality of sub-blocks in a row adjacent to the current block and a second plurality of sub-blocks in a column adjacent to the current block.

22. The apparatus of claim 20 , wherein the neighboring block is adjacent to the current block of video data.

23. The apparatus of claim 20 , wherein the neighboring block is not adjacent to the current block of video data.

24. The apparatus of claim 23 , wherein the affine model associated with the neighboring block is stored in a history table.

25. The apparatus of claim 24 , wherein the one or more processors are further configured to:

determine that the neighboring block is at a position that is outside of a distance constraint of the history table;

convert a coordinate of the neighboring block to be an adjusted coordinate located within the distance constraint of the history table; and

determine the affine model from a second affine model associated with a second block located at the adjusted coordinate.

26. The apparatus of claim 25 , wherein the one or more processors are further configured to:

determine the base motion vector field from a subset of sub-blocks of the second block based on the second block not being fully within the distance constraint of the history table.

27. The apparatus of claim 20 , wherein the one or more processors are further configured to:

determine the base motion vector field from a neighboring block coded using affine mode; and

determine the guidance motion vector field from motion information associated with a subset of a plurality of sub-blocks adjacent to the current block.

28. The apparatus of claim 27 , wherein the one or more processors are further configured to:

determine the subset of the plurality of sub-blocks adjacent to the current block based on a location of the neighboring block coded using the affine mode.

29. The apparatus of claim 27 , wherein the subset of the plurality of sub-blocks adjacent to the current block is one of a row of sub-blocks adjacent to the current block or a column of sub-blocks adjacent to the current block.

30. The apparatus of claim 27 , wherein the subset of the plurality of sub-blocks adjacent to the current block are within a distance constraint from the current block.

31. The apparatus of claim 27 , wherein to determine the guidance motion vector field from motion information associated with the subset of the plurality of sub-blocks adjacent to the current block, the one or more processors are further configured to:

sub-sample the plurality of sub-blocks adjacent to the current block.

32. The apparatus of claim 20 , wherein the guidance motion vector field is a sub-block temporal motion vector predictor field.

33. The apparatus of claim 20 , wherein the one or more processors are further configured to:

determine the base motion vector field by performing a linear regression process on two motion vector fields.

34. The apparatus of claim 20 , wherein the base motion vector field includes motion information from sub-blocks that are adjacent to the current block and motion information from sub-blocks that are not adjacent to the current block.

35. The apparatus of claim 20 , wherein to determine the refined affine model for the current block of video data, the one or more processors are further configured to:

determine the refined affine model using the base motion vector field, the guidance motion vector field, and at least one other motion vector field as inputs to a linear regression process.

36. The apparatus of claim 20 , wherein the one or more processors are further configured to:

scale motion vectors in one or more of the base motion vector field or the guidance motion vector field relative to a most frequently used reference picture index.

37. The apparatus of claim 20 , wherein the one or more processors are further configured to:

convert a uni-directional affine merge candidate of the affine merge candidates to a bi-directional affine merge candidate.

38. The apparatus of claim 20 , wherein the one or more processors are further configured to:

receive an index that specifies a particular affine merge candidate of the affine merge candidates;

determine a motion vector for the current block based on the particular affine merge candidate; and

decode the current block using the motion vector.

39. An apparatus for decoding video data, the apparatus comprising:

means for receiving a current block of video data to be decoded using an affine merge mode;

means for determining a base motion vector field from an affine model associated with a neighboring block;

means for determining a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

means for determining a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

means for determining affine merge candidates for the current block using the refined affine model; and

means for decoding the current block of video data using the affine merge candidates.

40. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors configured to decode video data to:

receive a current block of video data to be decoded using an affine merge mode;

determine a base motion vector field from an affine model associated with a neighboring block;

determine a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

determine a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

determine affine merge candidates for the current block using the refined affine model; and

decode the current block of video data using the affine merge candidates.

41. A method of encoding video data, the method comprising:

receiving a current block of video data to be encoded using an affine merge mode;

determining a base motion vector field from an affine model associated with a neighboring block;

determining a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

determining a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

determining affine merge candidates for the current block using the refined affine model; and

encoding the current block of video data using the affine merge candidates.

42. An apparatus configured to encode video data, the apparatus comprising:

a memory configured to store video data; and

one or more processors implemented in circuitry and in communication with the memory, the one or more processors configured to:

receive a current block of video data to be encoded using an affine merge mode;

determine a base motion vector field from an affine model associated with a neighboring block;

determine a guidance motion vector field from motion information associated with a plurality of sub-blocks adjacent to the current block;

determine a refined affine model for the current block of video data from a linear regression process using the base motion vector field and the guidance motion vector field as inputs to the linear regression process;

determine affine merge candidates for the current block using the refined affine model; and

encode the current block of video data using the affine merge candidates.

43. The method of claim 1 , further comprising:

determining that a coordinate of the neighboring block is beyond a distance relative to a coordinate of the current block;

converting the coordinate of the neighboring block to be an adjusted coordinate located within the distance relative to the coordinate of the current block; and

determining the affine model from a second affine model associated with a second block located at the adjusted coordinate.

44. The apparatus of claim 23 , wherein the one or more processors are further configured to:

determine that a coordinate of the neighboring block is beyond a distance relative to a coordinate of the current block;

convert the coordinate of the neighboring block to be an adjusted coordinate located within the distance relative to the coordinate of the current block; and

determine the affine model from a second affine model associated with a second block located at the adjusted coordinate.

45. The method of claim 41 , further comprising:

determining that a coordinate of the neighboring block is beyond a distance relative to a coordinate of the current block;

converting the coordinate of the neighboring block to be an adjusted coordinate located within the distance relative to the coordinate of the current block; and

determining the affine model from a second affine model associated with a second block located at the adjusted coordinate.

46. The apparatus of claim 42 , wherein the one or more processors are further configured to:

determine that a coordinate of the neighboring block is beyond a distance relative to a coordinate of the current block;

convert the coordinate of the neighboring block to be an adjusted coordinate located within the distance relative to the coordinate of the current block; and

determine the affine model from a second affine model associated with a second block located at the adjusted coordinate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: ZHANG, YAN; HUANG, HAN; SEREGIN, VADIM; COBAN, MUHAMMED ZEYD; KARCZEWICZ, MARTA
To: QUALCOMM INCORPORATED
Reel/Frame 063628/0223 →
Continuity (3)
Provisional Application 63366106 · Jun 9, 2022
Provisional Application 63362808 · Apr 11, 2022
Related Publication 20230328276A1 · Oct 12, 2023
References Cited (26)
US 20210243476A1 · Ko et al. · 2021 [cited by applicant]
US 20210266591A1 · Zhang et al. · 2021 [cited by applicant]
US 20210360277A1 · Jeong et al. · 2021 [cited by applicant]
US 20210385483A1 · Liu · 2021 [cited by examiner]
US 20220078488A1 · Leleannec et al. · 2022 [cited by applicant]
US 20220210462A1 · Luo · 2022 [cited by examiner]
WO 2020180704A1 · 2020 [cited by applicant]
Ghaznavi-Youvalari, R., “Regression-Based Motion Vector Field for Video Coding”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 5, (May 2021) (Year: 2021). [cited by examiner]
Ghaznavi-Youvalari R., et al., “CE2: Merge Mode with Regression-Based Motion Vector Field (Test 2.3.3)”, JVET-M0302, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marra… [cited by applicant]
International Search Report and Written Opinion—PCT/US2023/016231—ISA/EPO—Jun. 30, 2023. [cited by applicant]
Chen J., et al., “Algorithm Description for Versatile Video Coding and Test Model 9 (VTM 9)”, 130. MPEG Meeting, Mar. 20, 2020-Apr. 24, 2020, Alpbach, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. m53984,… [cited by applicant]
Chen W., et al., “AHG12: Non-Adjacent Spatial Neighbors for Affine Merge Mode”, JVET-X0151-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 24th Meeting, by teleconference, Oct. 6-15, 202… [cited by applicant]
Chen W., et al., “EE2-2.7, 2.8, 2.9: History-Parameter-Based Affine Model Inheritance and Non-Adjacent Affine Mode”, JVET-Z0139-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 26th Meeti… [cited by applicant]
Coban M., et al., “Algorithm Description of Enhanced Compression Model 2 (ECM 2)”, 23rd, MPEG Meeting, Jul. 7, 2021-Jul. 16, 2021, Online, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. M57745, JVET-W2025,… [cited by applicant]
Ghaznavi-Youvalari R., et al., “CE4-related: Merge Mode with Regression based Motion Vector Field (RMVF)”, JVET-L0171, Nokia Technologies, 2018, pp. 1-8. [cited by applicant]
Ghaznavi-Youvalari R., et al., “Regression-Based Motion Vector Field for Video Coding”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 5, May 2021, pp. 2034-2038. [cited by applicant]
Hu N., et al., “EE2-5: Adaptive filter Shape Switch and Using Samples before Deblocking Filter for Adaptive Loop Filter”, JVET-AA0095-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 27th… [cited by applicant]
ITU-T H.265: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, High Efficiency Video Coding, The International Telecommunication Union, Jun. 2019, 696 Pages. [cited by applicant]
ITU-T H.266: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, Versatile Video Coding, The International Telecommunication Union, Aug. 2020, 516 pages. [cited by applicant]
Karczewicz M., et al., “Common Test Conditions and Evaluation Procedures for Enhanced Compression Tool Testing”, JVET-Y2017-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 25th Meeting, … [cited by applicant]
Laroche G., et al., “EE2-2.5: ARMC Improvements”, JVET-AA0092-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 27th Meeting, by teleconference, Jul. 13-22, 2022, pp. 1-3. [cited by applicant]
Seregin V., et al., “Exploration Experiment on Enhanced Compression beyond VVC capability (EE2)”, JVET-Z2024-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 26th Meeting, by teleconferen… [cited by applicant]
Zhang K., et al., “An Improved Framework of Affine Motion Compensation in Video Coding”, IEEE Transactions on Image Processing (Early Access), Oct. 22, 2018, pp. 1-13. [cited by applicant]
Zhang K., et al., “EE2-3.12-Related: Extensions of History-Parameter-Based Affine Model Inheritance”, JVET-Y0161- 3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 25th Meeting, by teleconf… [cited by applicant]
Zhang Y., et al., “EE2-2.1: Regression Based Affine Candidate Derivation”, JVET-AA0107-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 27th Meeting, by teleconference, Jul. 13-22, 2022, … [cited by applicant]
Zhang Y., et al., “EE2-Related: Regression Based Affine Candidate Derivation”, JVET-Z0125-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 26th Meeting, by teleconference, Apr. 20-29, 202… [cited by applicant]