IP Library Granted Patent US 12,200,222
Granted Patent B2
US 12,200,222 · App. 18/521,810 · Granted Jan 14, 2025

Affine motion model derivation method

Inventors: Jiancong Luo (Skillman, NJ); Yuwen He (San Diego, CA); Wei Chen (San Diego, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/137H04N19/105H04N19/149H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,200,222
App. No.
18/521,810
Granted
Jan 14, 2025
Kind
B2
Abstract

Systems and methods are described for video coding using affine motion prediction. In an example method, motion vector gradients are determined from respective motion vectors of a plurality of neighboring sub-blocks neighboring a current block. An estimate of at least one affine parameter for the current block is determined based on the motion vector gradients. An affine motion model is determined based at least in part on the estimated affine parameter(s), and a prediction of the current block is generated using the affine motion model. The estimated parameter(s) may be used in the affine motion model itself. Alternatively, the estimated parameter(s) may be used in a prediction of the affine motion model. In some embodiments, only neighboring sub-blocks above and/or to the left of the current block are used in estimating the affine parameter(s).

Claims (40)

1. A video decoding method comprising:

for at least one current block in a video, determining at least one motion vector gradient from respective motion vectors of a plurality of neighboring sub-blocks neighboring the current block;

based at least in part on the motion vector gradient, selecting a motion model for prediction of the current block, wherein the selection is made from among a set of motion models including at least a four-parameter affine motion model and a six-parameter affine motion model; and

generating a prediction of the current block using the selected motion model.

2. The method of claim 1 , wherein the set of motion models further includes a translational motion model.

3. The method of claim 1 , wherein the selection of the motion model for prediction of the current block is based on a model fitting error.

4. The method of claim 1 , wherein selecting a motion model comprises, for each of a plurality of candidate motion models from among the set of motion models:

deriving estimated motion vectors for a plurality of the neighboring sub-blocks from the respective motion model, and determining a model fitting error based on differences between actual motion vectors of the plurality of neighboring sub-blocks and the corresponding estimated motion vectors of the plurality of neighboring sub-blocks;

wherein the selected motion model is selected to minimize the model fitting error.

5. The method of claim 4 , wherein the model fitting error is a weighted sum of differences between the actual motion vectors of the neighboring sub-blocks and the corresponding estimated motion vectors of the respective neighboring sub-blocks, each difference in the sum being weighed based on a distance between the current block and the respective neighboring sub-block.

6. A video decoder apparatus comprising one or more processors configured to perform at least:

for at least one current block in a video, determining at least one motion vector gradient from respective motion vectors of a plurality of neighboring sub-blocks neighboring the current block;

based at least in part on the motion vector gradient, selecting a motion model for prediction of the current block, wherein the selection is made from among a set of motion models including at least a four-parameter affine motion model and a six-parameter affine motion model; and

generating a prediction of the current block using the selected motion model.

7. The apparatus of claim 6 , wherein the set of motion models further includes a translational motion model.

8. The apparatus of claim 6 , wherein the selection of the motion model for prediction of the current block is based on a model fitting error.

9. The apparatus of claim 6 , wherein selecting a motion model comprises, for each of a plurality of candidate motion models from among the set of motion models:

deriving estimated motion vectors for a plurality of the neighboring sub-blocks from the respective motion model, and determining a model fitting error based on differences between actual motion vectors of the plurality of neighboring sub-blocks and the corresponding estimated motion vectors of the plurality of neighboring sub-blocks;

wherein the selected motion model is selected to minimize the model fitting error.

10. The apparatus of claim 9 , wherein the model fitting error is a weighted sum of differences between the actual motion vectors of the neighboring sub-blocks and the corresponding estimated motion vectors of the respective neighboring sub-blocks, each difference in the sum being weighed based on a distance between the current block and the respective neighboring sub-block.

11. A video encoding method comprising:

for at least one current block in a video, determining at least one motion vector gradient from respective motion vectors of a plurality of neighboring sub-blocks neighboring the current block;

based at least in part on the motion vector gradient, selecting a motion model for prediction of the current block, wherein the selection is made from among a set of motion models including at least a four-parameter affine motion model and a six-parameter affine motion model; and

generating a prediction of the current block using the selected motion model.

12. The method of claim 11 , wherein the set of motion models further includes a translational motion model.

13. The method of claim 11 , wherein the selection of the motion model for prediction of the current block is based on a model fitting error.

14. The method of claim 11 , wherein selecting a motion model comprises, for each of a plurality of candidate motion models from among the set of motion models:

deriving estimated motion vectors for a plurality of the neighboring sub-blocks from the respective motion model, and determining a model fitting error based on differences between actual motion vectors of the plurality of neighboring sub-blocks and the corresponding estimated motion vectors of the plurality of neighboring sub-blocks;

wherein the selected motion model is selected to minimize the model fitting error.

15. The method of claim 14 , wherein the model fitting error is a weighted sum of differences between the actual motion vectors of the neighboring sub-blocks and the corresponding estimated motion vectors of the respective neighboring sub-blocks, each difference in the sum being weighed based on a distance between the current block and the respective neighboring sub-block.

16. A video encoder apparatus comprising one or more processors configured to perform at least:

for at least one current block in a video, determining at least one motion vector gradient from respective motion vectors of a plurality of neighboring sub-blocks neighboring the current block;

based at least in part on the motion vector gradient, selecting a motion model for prediction of the current block, wherein the selection is made from among a set of motion models including at least a four-parameter affine motion model and a six-parameter affine motion model; and

generating a prediction of the current block using the selected motion model.

17. The apparatus of claim 16 , wherein the set of motion models further includes a translational motion model.

18. The apparatus of claim 16 , wherein the selection of the motion model for prediction of the current block is based on a model fitting error.

19. The apparatus of claim 16 , wherein selecting a motion model comprises, for each of a plurality of candidate motion models from among the set of motion models:

deriving estimated motion vectors for a plurality of the neighboring sub-blocks from the respective motion model, and determining a model fitting error based on differences between actual motion vectors of the plurality of neighboring sub-blocks and the corresponding estimated motion vectors of the plurality of neighboring sub-blocks;

wherein the selected motion model is selected to minimize the model fitting error.

20. The apparatus of claim 19 , wherein the model fitting error is a weighted sum of differences between the actual motion vectors of the neighboring sub-blocks and the corresponding estimated motion vectors of the respective neighboring sub-blocks, each difference in the sum being weighed based on a distance between the current block and the respective neighboring sub-block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: LUO, JIANCONG; HE, YUWEN; CHEN, WEI
To: VID SCALE, INC.
Reel/Frame 067359/0835 →
Continuity (3)
Continuation 17434974
Provisional Application 62814125 · Mar 5, 2019
Related Publication 20240107024A1 · Mar 28, 2024
References Cited (24)
US 20180077417A1 · Huang · 2018 [cited by examiner]
US 20180137632A1 · Takada · 2018 [cited by examiner]
US 20180192047A1 · Lv · 2018 [cited by examiner]
US 20180316929A1 · Li · 2018 [cited by examiner]
US 20190017811A1 · Watanabe · 2019 [cited by examiner]
US 20190028731A1 · Chuang · 2019 [cited by examiner]
US 20190045192A1 · Socek · 2019 [cited by examiner]
US 20190045214A1 · Ikai · 2019 [cited by examiner]
US 20190230350A1 · Chen · 2019 [cited by examiner]
US 20210368172A1 · Lim · 2021 [cited by examiner]
WO 2017157259A1 · 2017 [cited by applicant]
WO 2018067823A1 · 2018 [cited by applicant]
Bross, et al., “High Efficiency Video Coding (HEVC) Text Specification Draft 10 (for FDIS and Last Call)”. Joint Collaborative Team on Video Coding (JCT-VC), Document No. JCTVC-L1003, Jan. 2013, 310 pages. [cited by applicant]
International Telecommunication Union, “Advanced Video Coding for Generic Audiovisual Services”. Series H: Audiovisual and Multimedia System; Infrastructure of audiovisual services, Coding of moving video, ITU-T Recomme… [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority for PCT/US2020/020441, mailed May 12, 2020, 10 pages. [cited by applicant]
Ghaznavi-Youvalari, Ramin, et al., “CE4-related: Merge mode with Regression based Motion Vector Field (RMVF)”. Nokia, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-L0171, Oct. 3… [cited by applicant]
Ghaznavi-Youvalari, Ramin, et al. “CE2: Merge Mode with Regression-based Motion Vector Field (Test 2.3.3)”. Nokia, Joint Video Experts Team (JVET)of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-M0302, Jan. 9-18,… [cited by applicant]
He, Yuwen, et al., “CE4-Related: Affine Motion Estimation Improvements”. InterDigital Communications, Inc., Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-L0260-r2, Oct. 3-12, 20… [cited by applicant]
Chen, Wei, et al., “CE2-Related: Affine Motion Model Derivation Method for Affine Merge Mode”. InterDigital Communications, Inc., Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-N… [cited by applicant]
Wikipedia, “Sobel Filter”. Wikipedia web article available at: https://en.wikipedia.org/w/index.php?title=Sobel_operator&oldid=837616168, updated on Apr. 21, 2018, 9 pages. [cited by applicant]
International Preliminary Report on Patentability for PCT/US2020/20441 issued on Aug. 25, 2021, (7 pages). [cited by applicant]
Segall, Andrew, et al., “Joint Call for Proposals on Video Compression with Capability beyond HEVC”. Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-H1002, Oct. 18-24, 2017 (2… [cited by applicant]
Bross, Benjamin, et al., “Versatile Video Coding (Draft 2)”. Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-K1001-v6, Jul. 10-18, 2018 (139 pages). [cited by applicant]
SMPTE 421M, “VC-1 Compressed Video Bitstream Format and Decoding Process”. SMPTE Standard, 2006, (493 pages). [cited by applicant]