IP Library Granted Patent US 12,445,626
Granted Patent B2
US 12,445,626 · App. 18/755,674 · Granted Oct 14, 2025

Methods and apparatuses for prediction refinement with optical flow

Inventors: Xiaoyu Xiu (San Diego, CA); Yi-Wen Chen (San Diego, CA); Xianglin Wang (San Diego, CA); Shuiming Ye (San Diego, CA); Tsung-Chuan Ma (San Diego, CA); Hong-Jheng Jhu (San Diego, CA)
Assignee: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
H04N19/137H04N19/105H04N19/132H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,626
App. No.
18/755,674
Granted
Oct 14, 2025
Kind
B2
Abstract

Methods, apparatuses, and non-transitory computer-readable storage mediums are provided for coding a video signal. A method includes obtaining a first reference picture and a second reference picture associated with a video block; obtaining first prediction samples of the video block from the first reference picture; obtaining second prediction samples of the video block from the second reference picture; obtaining padded prediction samples, and obtaining horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the padded prediction samples; obtaining motion refinements for samples in the video block based on the horizontal and vertical gradient values; and obtaining bi-prediction samples of the video block based on the motion refinements.

Claims (47)

1. A method for coding a video signal, comprising:

obtaining a first reference picture and a second reference picture associated with a video block;

obtaining first prediction samples of the video block from the first reference picture;

obtaining second prediction samples of the video block from the second reference picture;

obtaining padded prediction samples, and obtaining horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the padded prediction samples, comprising:

deriving rows and columns of prediction samples outside the video block for the first prediction samples;

deriving rows and columns of prediction samples outside the video block for the second prediction samples; and

obtaining horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the derived rows and columns of prediction samples,

wherein bit-depths of the horizontal and vertical gradient values are controlled by performing a shift operation according to shift values, wherein the shift values comprise a first shift value for calculation of gradient values that are used in a prediction refinement with optical flow (PROF) process, and

wherein deriving the rows and columns of the prediction samples further comprises:

deriving a first part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a horizontal direction; and

deriving a second part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a vertical direction;

obtaining motion refinements for samples in the video block based on the horizontal and vertical gradient values; and

obtaining bi-prediction samples of the video block based on the motion refinements.

2. The method of claim 1 , wherein the rows and columns of prediction samples are copied from the corresponding integer reference samples.

3. A computing device for coding a video signal, comprising:

one or more processors;

a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors are configured to:

obtain a first reference picture and a second reference picture associated with a video block;

obtain first prediction samples of the video block from the first reference picture;

obtain second prediction samples of the video block from the second reference picture;

obtain padded prediction samples, and obtain horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the padded prediction samples, comprising:

derive rows and columns of prediction samples outside the video block for the first prediction samples;

derive rows and columns of prediction samples outside the video block for the second prediction samples; and

obtain horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the derived rows and columns of prediction samples,

wherein bit-depths of the horizontal and vertical gradient values are controlled by performing a shift operation according to shift values, wherein the shift values comprise a first shift value for calculation of gradient values that are used in a prediction refinement with optical flow (PROF) process, and

wherein derive the rows and columns of the prediction samples further comprises:

derive a first part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a horizontal direction; and

derive a second part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a vertical direction;

obtain motion refinements for samples in the video block based on the horizontal and vertical gradient values; and

obtain bi-prediction samples of the video block based on the motion refinements.

4. The computing device of claim 3 , wherein the rows and columns of prediction samples are copied from the corresponding integer reference samples.

5. A non-transitory computer readable storage medium storing a bitstream to be decoded by a coding method comprising:

obtaining a first reference picture and a second reference picture associated with a video block;

obtaining first prediction samples of the video block from the first reference picture;

obtaining second prediction samples of the video block from the second reference picture;

obtaining padded prediction samples, and obtaining horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the padded prediction samples, comprising:

deriving rows and columns of prediction samples outside the video block for the first prediction samples;

deriving rows and columns of prediction samples outside the video block for the second prediction samples; and

obtaining horizontal and vertical gradient values of the first prediction samples and the second prediction samples based on the derived rows and columns of prediction samples,

wherein bit-depths of the horizontal and vertical gradient values are controlled by performing a shift operation according to shift values, wherein the shift values comprise a first shift value for calculation of gradient values that are used in a prediction refinement with optical flow (PROF) process, and

wherein deriving the rows and columns of the prediction samples further comprises:

deriving a first part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a horizontal direction; and

deriving a second part of prediction samples from integer reference samples in the reference picture that are closest to a fractional sample position in a vertical direction;

obtaining motion refinements for samples in the video block based on the horizontal and vertical gradient values; and

obtaining bi-prediction samples of the video block based on the motion refinements.

6. The non-transitory computer readable storage medium of claim 5 , wherein the rows and columns of prediction samples are copied from the corresponding integer reference samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2024
From: XIU, XIAOYU; CHEN, YI-WEN; WANG, XIANGLIN; YE, SHUIMING; MA, TSUNG-CHUAN; JHU, HONG-JHENG
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 067929/0556 →
Continuity (4)
Continuation 17510328 · Oct 25, 2021
Continuation PCTUS2020030155 · Apr 27, 2020
Provisional Application 62838939 · Apr 25, 2019
Related Publication 20240348794A1 · Oct 17, 2024
References Cited (59)
US 11197019B2 · Abe · 2021 [cited by examiner]
US 20120031477A1 · Fogel et al. · 2012 [cited by applicant]
US 20140092968A1 · Guillemot et al. · 2014 [cited by applicant]
US 20150249825A1 · Kim et al. · 2015 [cited by applicant]
US 20160316217A1 · Hatakeyama · 2016 [cited by applicant]
US 20170150178A1 · Nam et al. · 2017 [cited by applicant]
US 20170238011A1 · Pettersson et al. · 2017 [cited by applicant]
US 20180192072A1 · Chen et al. · 2018 [cited by applicant]
US 20180199057A1 · Chuang et al. · 2018 [cited by applicant]
US 20180255298A1 · Lee · 2018 [cited by applicant]
US 20180262773A1 · Chuang et al. · 2018 [cited by applicant]
US 20180270500A1 · Li · 2018 [cited by applicant]
US 20180376147A1 · Iwamura et al. · 2018 [cited by applicant]
US 20180376166A1 · Chuang et al. · 2018 [cited by applicant]
US 20190045214A1 · Ikai et al. · 2019 [cited by applicant]
US 20200221122A1 · Ye et al. · 2020 [cited by applicant]
US 20200296405A1 · Huang et al. · 2020 [cited by applicant]
US 20200304827A1 · Abe · 2020 [cited by examiner]
US 20200382795A1 · Zhang et al. · 2020 [cited by applicant]
US 20220046249A1 · Xiu · 2022 [cited by examiner]
US 20220264115A1 · Li · 2022 [cited by examiner]
US 20230239479A1 · Jang · 2023 [cited by examiner]
US 20230269392A1 · Park · 2023 [cited by examiner]
CN 1856106A · 2006 [cited by applicant]
CN 101584210A · 2009 [cited by applicant]
CN 101795409A · 2010 [cited by applicant]
CN 106664416A · 2017 [cited by applicant]
CN 106664423A · 2017 [cited by applicant]
CN 107925775A · 2018 [cited by applicant]
CN 108496367A · 2018 [cited by applicant]
CN 109496430A · 2019 [cited by applicant]
EP 3357241A1 · 2018 [cited by applicant]
JP 2022523795A · 2022 [cited by applicant]
JP 7269371B2 · 2023 [cited by applicant]
JP 7559132B2 · 2024 [cited by applicant]
KR 20180107761A · 2018 [cited by applicant]
KR 20180119084A · 2018 [cited by applicant]
KR 20190024553A · 2019 [cited by applicant]
WO 2016056782A1 · 2016 [cited by applicant]
WO 2017036399A1 · 2017 [cited by applicant]
WO 2017195914A1 · 2017 [cited by applicant]
WO 2018131832A1 · 2018 [cited by applicant]
WO 2018113658A1 · 2018 [cited by applicant]
WO 2018166357A1 · 2018 [cited by applicant]
WO 2018237303A1 · 2018 [cited by applicant]
WO 2019072371A1 · 2019 [cited by applicant]
WO 2019010156A1 · 2019 [cited by applicant]
WO 2020211867A1 · 2020 [cited by applicant]
Bross, Benjamin et al., Versatile Video Coding (Draft 5), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WVG 11, 14th Meeting: Geneva, CH, [Document: VET-N1001-v101, Mar. 27, 2019, (407p). [cited by applicant]
Yan-fei Shen et al. “High Efficiency Video Coding”, Chinese Journal of Computer, vol. 36 No. 1 dated on Nov. 2013 (16p). [cited by applicant]
Yu-jiao Wang, “Study On Motion Blurred Video Images' Restoration”, Xia'n University of Technology, Student No. 2140920003, dated on Jun. 2017 (58p). [cited by applicant]
Luo, et al., CE2-related: Prediction refinement whti optical flow for affine mode', Joint Video Experts Team (JVET) of ITU-T SG 61 WP 3and ISO/IEC JC 1/SC 29/WG 1, JVET-N0236-15, 14th Meeting: Geneva, CH, Mar. 19-27, 20… [cited by applicant]
Yang, Haito, “Description of Core Experiment 4(CE4): Inter prediction”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/EC JTC 1/SC 29/WG11, JVET-N1024-v3, 14th Meeting: Geneva, CH, Mar. 19-27, 2019, (11p). [cited by applicant]
Supplementary EP Search Report of EP Application No. 20795945.3 dated Nov. 25, 2022, (4p). [cited by applicant]
Bross, Benjamin et al., “Versatile Video Coding (Draft 4)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/EC JTC 1/SC 29/WG11, JVET-M1001-v-7, 13th Meeting: Marrakech, MA Jan. 9-18, 2019, (299p). [cited by applicant]
International Search Report of International Application No. PCT/US2020/030155 dated Jul. 30, 2020, (3p). [cited by applicant]
Luo, Jiancong (Daniel) et al., CE2-Related: Prediction refinement with optical flow for affine mode, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 1, 14th Meeting: Geneva, CH [Document: … [cited by applicant]
Xiaoyu Xiu et al., “CE4-related: Harmonization of BDOF and PROF”, oint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, JVET-00593, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019, (5p). [cited by applicant]
EP Office Action of European Patent Application No. 20795945.3 dated Mar. 4, 2025, (7p). [cited by applicant]