IP Library Granted Patent US 12,238,279
Granted Patent B2
US 12,238,279 · App. 17/611,699 · Granted Feb 25, 2025

Bi-prediction refinement in affine with optical flow

Inventors: Ya Chen (Cesson-Sevigne, FR); Franck Galpin (Cesson-Sevigne, FR); Fabrice Leleannec (Cesson-Sevigne, FR); Antoine Robert (Cesson-Sevigne, FR)
Assignee: InterDigital CE Patent Holdings, SAS
H04N19/107H04N19/105H04N19/176H04N19/513H04N19/577H04N19/523H04N19/563
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,279
App. No.
17/611,699
Granted
Feb 25, 2025
Kind
B2
Abstract

PROF (Prediction Refinement with Optical Flow) and BDOF (Bi-directional Optical Flow) use the optical flow model to provide sample-wise motion compensation refinement. In one implementation, to calculate the spatial gradients in BDOF, an extended region is generated around a block to be coded. To simplify the computation, BDOF may use the same padding operations as in PROF, namely, a padding sample is copied from the nearest integer neighbor in the reference picture. BDOF and PROF can also be used together when sub-block based affine motion-compensated prediction is used. BDOF can be applied after PROF generates new prediction blocks, or BDOF can use the spatial gradients in the original prediction blocks to re-use the computations with PROF. In addition, BDOF and PROF operations can be combined, and the prediction adjustment will be directly based on the spatial gradients in the original prediction blocks.

Claims (52)

1. A method, comprising:

accessing a block to be decoded in affine mode, said block including a plurality of sub-blocks;

for one sub-block of said plurality of sub-blocks, obtaining a first and second prediction block, from a first and second reference picture, respectively, using sub-block based affine motion-compensated prediction, said one block in a bi-prediction mode;

obtaining motion refinement for said one sub-block, based on said first and second prediction blocks and spatial gradients in said first and second prediction blocks, according to a temporal optical flow model and a spatial optical flow model, wherein said motion refinement for said one sub-block is based on a linear combination of a motion difference according to said spatial optical flow model and an initial motion refinement for said one sub-block according to said temporal optical flow model;

obtaining a prediction adjustment for said one sub-block, based on said motion refinement and said spatial gradients; and

decoding said one sub-block based on said prediction adjustment, said first prediction block and said second prediction block.

2. The method of claim 1 , wherein said obtaining a prediction adjustment comprises:

scaling a horizontal gradient and a vertical gradient for said first prediction block by a horizontal motion refinement and a vertical motion refinement for said first prediction block, respectively; and

scaling a horizontal gradient and a vertical gradient for said second prediction block by a horizontal motion refinement and a vertical motion refinement for said second prediction block, respectively.

3. The method of claim 2 , further comprising:

obtaining a motion difference between a motion vector at a sample level and a motion vector for said first prediction block, for each sample in said first prediction block,

wherein said horizontal motion refinement and said vertical motion refinement for said first prediction block are obtained based on said motion difference.

4. The method of claim 1 , wherein said first and second prediction blocks and spatial gradients in said first and second prediction blocks are directly used for said motion refinement and said prediction adjustment for said one sub-block.

5. An apparatus, comprising one or more processors, wherein said one or more processors are configured to:

access a block to be decoded in affine mode, said block including a plurality of sub-blocks;

for one sub-block of said plurality of sub-blocks, obtain a first and second prediction block, from a first and second reference picture, respectively, using sub-block based affine motion-compensated prediction, said one block in a bi-prediction mode;

obtain motion refinement for said one sub-block, based on said first and second prediction blocks and spatial gradients in said first and second prediction blocks, according to a temporal optical flow model and a spatial optical flow model, wherein said motion refinement for said one sub-block is based on a linear combination of a motion difference according to said spatial optical flow model and an initial motion refinement for said one sub-block according to said temporal optical flow model;

obtain a prediction adjustment for said one sub-block, based on said motion refinement and said spatial gradients; and

decode said one sub-block based on said prediction adjustment, said first prediction block and said second prediction block.

6. The apparatus of claim 5 , wherein said one or more processors are further configured to:

scale a horizontal gradient and a vertical gradient for said first prediction block by a horizontal motion refinement and a vertical motion refinement for said first prediction block, respectively; and

scale a horizontal gradient and a vertical gradient for said second prediction block by a horizontal motion refinement and a vertical motion refinement for said second prediction block, respectively.

7. The apparatus of claim 6 , wherein said one or more processors are further configured to:

obtain a motion difference between a motion vector at a sample level and a motion vector for said first prediction block, for each sample in said first prediction block,

wherein said horizontal motion refinement and said vertical motion refinement for said first prediction block are obtained based on said motion difference.

8. The apparatus of claim 5 , wherein said first and second prediction blocks and spatial gradients in said first and second prediction blocks are directly used for said motion refinement and said prediction adjustment for said one sub-block.

9. A method, comprising:

accessing a block to be encoded in affine mode, said block including a plurality of sub-blocks;

for one sub-block of said plurality of sub-blocks, obtaining a first and second prediction block, from a first and second reference picture, respectively, using sub-block based affine motion-compensated prediction, said one block in a bi-prediction mode;

obtaining motion refinement for said one sub-block, based on said first and second prediction blocks and spatial gradients in said first and second prediction blocks, according to a temporal optical flow model and a spatial optical flow model, wherein said motion refinement for said one sub-block is based on a linear combination of a motion difference according to said spatial optical flow model and an initial motion refinement for said one sub-block according to said temporal optical flow model;

obtaining a prediction adjustment for said one sub-block, based on said motion refinement and said spatial gradients; and

encoding said one sub-block responsive to said prediction adjustment, said first prediction block and said second prediction block.

10. The method of claim 9 , wherein said obtaining a prediction adjustment comprises:

scaling a horizontal gradient and a vertical gradient for said first prediction block by a horizontal motion refinement and a vertical motion refinement for said first prediction block, respectively; and

scaling a horizontal gradient and a vertical gradient for said second prediction block by a horizontal motion refinement and a vertical motion refinement for said second prediction block, respectively.

11. The method of claim 10 , further comprising:

obtaining a motion difference between a motion vector at a sample level and a motion vector for said first prediction block, for each sample in said first prediction block,

wherein said horizontal motion refinement and said vertical motion refinement for said first prediction block is obtained based on said motion difference.

12. The method of claim 9 , wherein said first and second prediction blocks and spatial gradients in said first and second prediction blocks are directly used for said motion refinement and said prediction adjustment for said one sub-block.

13. An apparatus, comprising one or more processors, wherein said one or more processors are configured to:

access a block to be encoded in affine mode, said block including a plurality of sub-blocks;

for one sub-block of said plurality of sub-blocks, obtain a first and second prediction block, from a first and second reference picture, respectively, using sub-block based affine motion-compensated prediction, said one block in a bi-prediction mode;

obtain motion refinement for said one sub-block, based on said first and second prediction blocks and spatial gradients in said first and second prediction blocks, according to a temporal optical flow model and a spatial optical flow model, wherein said motion refinement for said one sub-block is based on a linear combination of a motion difference according to said spatial optical flow model and an initial motion refinement for said one sub-block according to said temporal optical flow model;

obtain a prediction adjustment for said one sub-block, based on said motion refinement and said spatial gradients; and

encode said one sub-block responsive to said prediction adjustment, said first prediction block and said second prediction block.

14. The apparatus of claim 13 , wherein said one or more processors are further configured to:

scale a horizontal gradient and a vertical gradient for said first prediction block by a horizontal motion refinement and a vertical motion refinement for said first prediction block, respectively; and

scale a horizontal gradient and a vertical gradient for said second prediction block by a horizontal motion refinement and a vertical motion refinement for said second prediction block, respectively.

15. The apparatus of claim 14 , wherein said one or more processors are further configured to:

obtain a motion difference between a motion vector at a sample level and a motion vector for said first prediction block, for each sample in said first prediction block,

wherein said horizontal motion refinement and said vertical motion refinement for said first prediction block is obtained based on said motion difference.

16. The apparatus of claim 13 , wherein said first and second prediction blocks and spatial gradients in said first and second prediction blocks are directly used for said motion refinement and said prediction adjustment for said one sub-block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2023
From: INTERDIGITAL VC HOLDINGS FRANCE, SAS
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 064396/0118 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2021
From: CHEN, YA; GALPIN, FRANCK; LELEANNEC, FABRICE; ROBERT, ANTOINE
To: INTERDIGITAL VC HOLDING FRANCE, SAS
Reel/Frame 058137/0725 →
Priority Claims (1)
EP 19305885 · Jul 1, 2019 · regional
Continuity (1)
Related Publication 20220264146A1 · Aug 18, 2022
References Cited (58)
US 20200169748A1 · Chen · 2020 [cited by examiner]
US 20200221117A1 · Liu · 2020 [cited by examiner]
US 20200351495A1 · Li · 2020 [cited by examiner]
US 20200366928A1 · Liu · 2020 [cited by examiner]
US 20200382795A1 · Zhang · 2020 [cited by examiner]
US 20200413082A1 · Li · 2020 [cited by examiner]
US 20210051339A1 · Liu · 2021 [cited by examiner]
US 20210144400A1 · Liu · 2021 [cited by examiner]
US 20210226585A1 · Khlat · 2021 [cited by examiner]
US 20210227211A1 · Liu · 2021 [cited by examiner]
US 20210227250A1 · Liu · 2021 [cited by examiner]
US 20210274213A1 · Xiu · 2021 [cited by examiner]
US 20210329257A1 · Sethuraman · 2021 [cited by examiner]
US 20210367203A1 · Pi · 2021 [cited by examiner]
US 20210368198A1 · Zhang · 2021 [cited by examiner]
US 20210368199A1 · Zhang · 2021 [cited by examiner]
US 20210368203A1 · Zhang · 2021 [cited by examiner]
US 20210385481A1 · Liu · 2021 [cited by examiner]
US 20210385482A1 · Liu · 2021 [cited by examiner]
US 20220046249A1 · Xiu · 2022 [cited by examiner]
US 20220060690A1 · Sethuraman · 2022 [cited by examiner]
US 20220094974A1 · Toma · 2022 [cited by examiner]
US 20220103827A1 · Liu · 2022 [cited by examiner]
US 20220116655A1 · Xiu · 2022 [cited by examiner]
US 20220182658A1 · Xiu · 2022 [cited by examiner]
US 20220210462A1 · Luo · 2022 [cited by examiner]
US 20220232244A1 · Li · 2022 [cited by examiner]
US 20220286688A1 · Chen · 2022 [cited by examiner]
US 20240121431A1 · Xiu · 2024 [cited by examiner]
CA 3131311A1 · 2020 [cited by examiner]
CA 3144099A1 · 2020 [cited by examiner]
CA 3144797A1 · 2020 [cited by examiner]
CA 3145056A1 · 2020 [cited by examiner]
TW 202025766A · 2020 [cited by examiner]
WO WO2020003260A1 · 2020 [cited by examiner]
WO WO2020016857A1 · 2020 [cited by examiner]
WO WO2020058890A1 · 2020 [cited by examiner]
WO WO2020070729A1 · 2020 [cited by examiner]
WO WO2020084476A1 · 2020 [cited by examiner]
WO WO2020251419A2 · 2020 [cited by examiner]
Machine translation of TW-202025766-A (Year: 2020). [cited by examiner]
A. Alexander & A. Elena, “Bi-directional optical flow for future video codec”, 2016 Data Compression Conf. 83-90 (Apr. 2016) (Year: 2016). [cited by examiner]
A. Alshin, E. Alshina, & T. Lee, “Bi-Directional Optical Flow for Improving Motion Compensation”, 28 Picture Coding Symp. 422-425 (Dec. 2010) (Year: 2010). [cited by examiner]
H. Gao, S. Esenlik, Z. Zhao, E. Steinbach, & J. Chen, “Decoder Side Motion Vector Refinement for Versatile Video Coding”, presented at 21 IEEE Workshop on Multimedia Signal Processing (Sep. 2019) (Year: 2019). [cited by examiner]
S. Kamp, M. Evertz, & M. Wien, “Decoder Side Motion Vector Derivation for Inter Frame Video Coding”, 15 IEEE Int'l Conf. on Image Processing 1120-23 (Oct. 2008) (Year: 2008). [cited by examiner]
S. Kamp & M. Wien, “Decoder-side Motion Vector Derivation for Block-Based Video Coding”, 22 IEEE Transactions on Circuits & Sys. for Video Tech 1732-1745 (Dec. 2012) (Year: 2012). [cited by examiner]
Bross et al., “Versatile Video Coding (Draft 3)”, Document: JVET-L1001-v9, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, 3-12, pp. 1-233, Oct. 12, 2018. [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 5)”, Document: JVET-N1001-v8, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Gevenva, CH, pp. 1-397, Mar. 19-27, 2019. [cited by applicant]
Luo et al., “CE2-related: Prediction refinement with optical flow for affine mode”, Document: JVET-N0236-r5, J Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH,… [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 5) ”, Document: JVET-N1001-v5, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH, pp. 1-370, Mar. 19-27, 2019. [cited by applicant]
Chen J et al: “Algorithm description for Versatile Video Coding and Test Model 5 (VTM 5)”, Document: JVET-N1002-v2 of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH, pp. 1-76, Mar. 19-27, 2019. [cited by applicant]
Anonymous, “High efficiency video coding”, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual servic… [cited by applicant]
Xiu et al, “CE4-related: Prediction sample padding unification for BDOF and PROF”, Document: JVET-O0594, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, p… [cited by applicant]
Lee et al, “CE9-related; A simple gradient calculation at the CU boundaries for BDOF”, Document: JVET-M0241, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA… [cited by applicant]
Galpin et al., “CE9-related: combination of PROF for affine and BDOF-BWA”, JVET of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, document JVET-N0757, 2019, pp. 1-2. [cited by applicant]
Xiu (Kwai Inc.) et al: “CE4-related: Harmonization of BDOF and PROF”, JVET of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting, document JVET-00593, 2019, pp. 1-5. [cited by applicant]
Li (Panasonic Corporation) et al: “CE4-related: Alignment of BDOF refinement process with PROF”, JVET of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting, document JVET-00123-v3, 2019, pp. 1-3. [cited by applicant]
Galpin (Technicolor) et al: “CE9-related: BDOF-BWA unification”, JVET of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, document JVET-N0239, 2019, pp. 1-3. [cited by applicant]
Cited By (1)
US 12,598,291