IP Library Granted Patent US 12,341,969
Granted Patent B2
US 12,341,969 · App. 17/442,536 · Granted Jun 24, 2025

Methods and apparatus for prediction refinement for decoder side motion vector refinement with optical flow

Inventors: Wei Chen (San Diego, CA); Yuwen He (San Diego, CA); Jiancong Luo (Skillman, NJ)
Assignee: InterDigital VC Holdings, Inc.
H04N19/137H04N19/105H04N19/176H04N19/513
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,341,969
App. No.
17/442,536
Granted
Jun 24, 2025
Kind
B2
Abstract

Methods, devices, apparatus, systems, architectures and interfaces to improve motion vector (MV) refinement based sub-block (SB) level motion compensated prediction are provided. A decoding method includes receiving a bitstream of encoded video data, the bitstream including at least one block of video data including a plurality of SBs; performing a MV derivation, including a decoder based MV (DMVR) process, for at least one SB in the block to generate a refined MV for each SB; performing SB based motion compensation on the at least one sub-block to generate a SB based prediction within each SB; obtaining a spatial gradient for the prediction within each SB; determining a MV offset for each pixel in each SB; obtaining an intensity change in each SB based on the spatial gradients and MV offsets via an optical flow equation; and refining the prediction within each SB based on the obtained intensity changes.

Claims (70)

1. A method for video decoding, comprising:

receiving a block of video data, wherein the block of video data includes a plurality of sub-blocks; and

for a sub-block of the plurality of sub-blocks, refining a prediction of the sub-block by:

deriving motion vectors, each indicating a refinement motion of a prediction of a respective one of the plurality of sub-blocks,

computing an affine motion model based on the derived motion vectors,

determining, for pixels of the prediction of the sub-block, respective motion vector offsets using the affine motion model,

obtaining an intensity change for the pixels based on respective spatial gradients and the respective motion vector offsets via an optical flow equation, and

refining the pixels based on the obtained intensity change.

2. The method of claim 1 , wherein the derived motion vectors are obtained by a decoder-based motion vector refinement process.

3. The method of claim 1 , wherein a model error is estimated based on the derived motion vectors, and if the estimated model error exceeds a threshold the computing, the determining, the obtaining, and the refining for the sub-block are skipped.

4. The method of claim 1 , wherein the plurality of sub-blocks includes sub-blocks neighboring the sub-block and wherein the computing of the affine motion model is based on motion vectors derived for the neighboring sub-blocks, and further comprising:

estimating a model error of the affine motion model using the motion vector derived for the sub-block; and

if the estimated model error exceeds a threshold, skipping the determining, the obtaining, and the refining for the sub-block.

5. The method of claim 1 , wherein the affine motion model computed for the sub-block is used for refining a prediction of another sub-block of the video block.

6. The method of claim 1 , wherein the affine motion model is a four-parameter affine motion model for one or more sub-blocks of the plurality of sub-blocks, and wherein the one or more sub-blocks are located at a left boundary and/or a top boundary of the block of video data.

7. The method of claim 1 , wherein the computing of the affine motion model further comprises:

computing multiple models based on the derived motion vectors;

estimating respective model errors for the multiple models; and

selecting as the affine motion model one of the multiple models having the least model error.

8. The method of claim 1 , wherein the computing of the affine motion model further comprises:

computing multiple models based on the derived motion vectors, wherein the multiple models include one or more of a six-parameter affine motion model and a four-parameter affine motion model.

9. The method of claim 1 , wherein the refining the pixels based on the obtained intensity change comprises:

weighting the obtained intensity change by a weighting factor, and

adding the weighted intensity change to the pixels.

10. The method of claim 9 , further comprising:

receiving a weight index indicating, at the video block level or at a picture level, the weighting factor for weighting the obtained intensity change.

11. A method for video encoding, comprising:

receiving a block of video data, wherein the block of video data includes a plurality of sub-blocks; and

for a sub-block of the plurality of sub-blocks, refining a prediction of the sub-block by:

deriving motion vectors, each indicating a refinement motion of a prediction of a respective one of the plurality of sub-blocks,

computing an affine motion model based on the derived motion vectors,

determining, for pixels of the prediction of the sub-block, respective motion vector offsets using the affine motion model,

obtaining an intensity change for the pixels based on respective spatial gradients and the respective motion vector offsets via an optical flow equation, and

refining the pixels based on the obtained intensity change.

12. An apparatus for video decoding, comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the apparatus to:

receive a block of video data, wherein the block of video data includes a plurality of sub-blocks, and

for a sub-block of the plurality of sub-blocks, refine a prediction of the sub-block by:

deriving motion vectors, each indicating a refinement motion of a prediction of a respective one of the plurality of sub-blocks,

computing an affine motion model based on the derived motion vectors,

determining, for pixels of the prediction of the sub-block, respective motion vector offsets using the affine motion model,

obtaining an intensity change for the pixels based on respective spatial gradients and the respective motion vector offsets via an optical flow equation, and

refining the pixels based on the obtained intensity change.

13. The apparatus of claim 12 , wherein the derived motion vectors are obtained by a decoder-based motion vector refinement process.

14. The apparatus of claim 12 , wherein the plurality of sub-blocks includes sub-blocks neighboring the sub-block, wherein the computing of the affine motion model is based on motion vectors derived for the neighboring sub-blocks, and wherein the instructions further cause the apparatus to:

estimate a model error of the affine motion model using the motion vector derived for the sub-block; and

if the estimated model error exceeds a threshold, skip the determining, the obtaining, and the refining for the sub-block.

15. The apparatus of claim 12 , wherein the affine motion model computed for the sub-block is used for refining a prediction of another sub-block of the video block.

16. The apparatus of claim 12 , wherein the computing of the affine motion model further comprises:

computing multiple models based on the derived motion vectors;

estimating respective model errors for the multiple models; and

selecting as the affine motion model one of the multiple models having the least model error.

17. The apparatus of claim 12 , wherein the computing of the affine motion model further comprises:

computing multiple models based on the derived motion vectors, wherein the multiple models include one or more of a six-parameter affine motion model and a four-parameter affine motion model.

18. The apparatus of claim 12 , wherein the refining the pixels based on the obtained intensity change comprises:

weighting the obtained intensity change by a weighting factor, and

adding the weighted intensity change to the pixels.

19. The apparatus of claim 18 , further comprising:

receiving a weight index indicating, at the video block level or at a picture level, the weighting factor for weighting the obtained intensity change.

20. An apparatus for video encoding, comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the apparatus to:

receive a block of video data, wherein the block of video data includes a plurality of sub-blocks, and

for a sub-block of the plurality of sub-blocks, refine a prediction of the sub-block by:

deriving motion vectors, each indicating a refinement motion of a prediction of a respective one of the plurality of sub-blocks,

computing an affine motion model based on the derived motion vectors,

determining, for pixels of the prediction of the sub-block, respective motion vector offsets using the affine motion model,

obtaining an intensity change for the pixels based on respective spatial gradients and the respective motion vector offsets via an optical flow equation, and

refining the pixels based on the obtained intensity change.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2021
From: CHEN, WEI; HE, YUWEN; LUO, JIANCONG
To: VID SCALE, INC.
Reel/Frame 057754/0532 →
Continuity (2)
Provisional Application 62823935 · Mar 26, 2019
Related Publication 20220191502A1 · Jun 16, 2022
References Cited (34)
US 11695950B2 · Luo · 2023 [cited by examiner]
US 12143626B2 · Luo · 2024 [cited by examiner]
US 20100246680A1 · Tian · 2010 [cited by examiner]
US 20180262773A1 · Chuang et al. · 2018 [cited by applicant]
US 20180270500A1 · Li et al. · 2018 [cited by applicant]
US 20200296405A1 · Huang · 2020 [cited by examiner]
CN 109155855A · 2019 [cited by applicant]
EP 3422720A1 · 2019 [cited by applicant]
RU 2671307C1 · 2018 [cited by applicant]
RU 2696551C1 · 2019 [cited by applicant]
TW 201907728A · 2019 [cited by applicant]
WO WO2019002215A1 · 2019 [cited by applicant]
WO WO2020191034A1 · 2020 [cited by applicant]
Chen et al., Algorithm description for Versatile Video Coding and Test Model 4 (VTM4), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG11, Document: JVET-M 1002-v2; 13th Meeting: Marrakech,… [cited by examiner]
Luo et al., “CE2-related: Prediction refinement with optical flow for affine mode”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-N0236-r1, 14th Meeting: Geneva, Switz… [cited by applicant]
SMPTE 421M, “VC-1 Compressed Video Bitstream Format and Decoding Process”,. SMPTE Standard, Apr. 2006, 493 pages. [cited by applicant]
“BMS-2.0 Reference Software”, available at <https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_BMS/tags/BMS-2.1rc1>, 1 page. [cited by applicant]
International Telecommunication Union (ITU), Advanced Video Coding for Generic Audiovisual Services, Series H: Audiovisual and Multimedia Systems; Infrastructure of Audiovisual Services—Coding of Moving Video; ITU-T Tec… [cited by applicant]
Luo J., et al.,: “CE4-2.1: Prediction refinement with optical flow for affine mode”, 15. JVET Meeting; Jul. 3-12, 2019; Gothenburg; (JVET of ISO/IEC JTC1/SC29/WG11 and ITU-T SG.16), No. JVET-00070 Jun. 26, 2019, XP03021… [cited by applicant]
Chen W., et al., “Non-CE9: Block Boundary Prediction Refinement with Optical Flow for DMVR”, 127. MPEG Meeting; Jul. 8-12, 2019; Gothenburg; (MPEG or ISO/IEC JTC1/SC29/WG11), No. m48720 Jun. 26, 2019, XP030222192, Retri… [cited by applicant]
Suhring, Karsten et al., “H.264/14496-10 AVC Reference Software Manual (Revised for JM 19.0)”, Document JVT-AE010, Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG (ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q.6), 31st Me… [cited by applicant]
Chen et al., Algorithm description for Versatile Video Coding and Test Model 4 (VTM 4), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-M1002-v2; 13th Meeting: Marrakech… [cited by applicant]
“HM Reference Software HM-16.9”, available at <https://hevc.hhi.fraunhofer.de/svn/svn_HEVCSoftware>, Mar. 2016, 2 pages. [cited by applicant]
Sobel filter, https://en.wikipedia.org/wiki/Sobel_operator, 7 pages. [cited by applicant]
Segall et al., “Joint Call for Proposals on Video Compression with Capability beyond HEVC”, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Doc. JVET-H1002, 8th Meeting: Macao, CN,… [cited by applicant]
“JEM-7.0 Reference Software”, available at <https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/HM-16.6-JEM-7.0>, 1 page. [cited by applicant]
Bross et al, Versatile Video Coding (Draft 2), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-K1001-v1, 11th Meeting: Ljubljana, Slovenia, Jul. 10, 2018, 41 pages. [cited by applicant]
“Test Model 4 of Versatile Video Coding (VTM 3)”, 125. MPEG Meeting; Jan. 14-18, 2019; Marrakech; (MPEG or ISO/IEC JTC1/SC29/WG11), No. n18275 Mar. 24, 2019, XP030212829, Retrieved from the Internet: URL:http://phenix.i… [cited by applicant]
Chen et al., “Description of SDR, HDR and 360° video coding technology proposal by Huawei, GoPro, HiSilicon, and Samsung”, JVET-J0025, buJoint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG… [cited by applicant]
Bross et al., “High Efficiency Video Coding (HEVC) text specification draft 10”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JCTVC-L1003, Feb. 2012, 321… [cited by applicant]
“VTM-2.0.1 reference software”, available at https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/tags/VTM-2.0.1. [cited by applicant]
Huang et al., CE2-related: Alignment of affine control-point motion vector and subblock motion vector, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-M0110, 13th Meeting: Marrake… [cited by applicant]
Li et al., An Efficient Four-Parameter Affine Motion Model for Video Coding, IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, Issue: 8, Aug. 2018. [cited by applicant]
Philippe Hanhart and Yuwen He, Non-CE2: Motion vector clipping in affine sub-block motion vector derivation, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-M0145-v1, 13th Meeting… [cited by applicant]