IP Library › Granted Patent US 12,464,125
Granted Patent B2
US 12,464,125 · App. 18/373,113 · Granted Nov 4, 2025

Method and apparatus for video coding using deep learning based in-loop filter for inter prediction

Inventors: Je Won Kang (Seoul, KR); Na Young Kim (Seoul, KR); Jung Kyung Lee (Seoul, KR); Seung Wook Park (Yongin-si, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION; EWHA UNIVERSITY—INDUSTRY COLLABORATION FOUNDATION
H04N19/124G06T5/20G06T5/60G06T5/70G06V10/771H04N19/132H04N19/159
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,464,125
App. No.
18/373,113
Granted
Nov 4, 2025
Kind
B2
Abstract

A method and an apparatus for video coding using a deep learning-based in-loop filter for inter-prediction are disclosed. The video coding method and the apparatus utilize a deep learning-based in-loop filter for inter-prediction of a predictive frame (P-frame) and a bipredictive frame (B-frame) in order to mitigate various levels of image distortion according to a QP (quantization parameter) value present in the P-frame and the B-frame.

Claims (42)

1 . A method for filtering an image area performed by a video decoding device, the method comprising:

obtaining an image area having been reconstructed and a quantization parameter of the image area;

generating an embedding vector based on the quantization parameter and a prediction type of the image area; and

generating a filtered image area of the image area based on the embedding vector by using a denoising model that is based on deep learning,

wherein the prediction type of the image area indicates one of an intra prediction type in which the image area is predicted independently, a predictive type in which the image area is predicted based on a reference image area in a single direction, or a bi-predictive type in which the image area is predicted based on at least one reference image area in bi-directions.

2 . The method of claim 1 , wherein the image area is a predictive frame (P-frame) or a bipredictive frame (B-frame) reconstructed according to an inter-prediction.

3 . The method of claim 1 , wherein generating the embedding vector includes:

generating the embedding vector using embedding function including an embedding layer and a plurality of fully-connected layers.

4 . The method of claim 3 , wherein the embedding function takes as input all or part of the quantization parameter, a Lagrange multiplier for calculating rate distortion, a temporal layer of the image area, a type of the image area, or any combination thereof.

5 . The method of claim 1 , wherein

the denoising model includes a cascaded structure of residual blocks (RBs) and convolutional layers and uses the cascaded structure to generate the filtered image area, and

each RB is a convolutional block having a skip path between an input and an output.

6 . The method of claim 5 , wherein generating the filtered image area includes multiplying a feature generated by a preset convolutional layer among the convolutional layers by an absolute value of the embedding vector.

7 . The method of claim 1 , wherein the denoising model includes:

a U-net that is a deep learning model configured to generate an offset of a kernel from the image area;

a sampler configured to sample the image area by using the offset;

convolutional layers configured to generate a calibrated kernel from the image area, an output feature map of the U-net, and a sampled image area; and

an output convolutional layer configured to apply convolution to the sampled image area by using the calibrated kernel to generate the filtered image area.

8 . The method of claim 7 , wherein generating the filtered image area includes multiplying the calibrated kernel by an absolute value of the embedding vector.

9 . The method of claim 1 , wherein

the denoising model further includes combinatorial convolutional layers,

the denoising model generates residual signals between the image area and the filtered image area by using an absolute value of the embedding vector and the combinatorial convolutional layers, and

the denoising model sums the residual signals and the filtered image area.

10 . A method performed by a video encoding device for filtering an image area, the method comprising:

obtaining an image area having been reconstructed and a quantization parameter of the image area;

generating an embedding vector based on the quantization parameter and a prediction type of the image area; and

generating a filtered image area based on the embedding vector by using a denoising model that is based on deep learning,

wherein the prediction type of the image area indicates one of an intra prediction type in which the image area is predicted independently, a predictive type in which the image area is predicted based on a reference image area in a single direction, or a bi-predictive type in which the image area is predicted based on at least one reference image area in bi-directions.

11 . The method of claim 10 , wherein obtaining the image area and the quantization parameter comprises:

obtaining as the image area a predictive frame (P-frame) or a bipredictive frame (B-frame) reconstructed according to an inter-prediction.

12 . The method of claim 10 , wherein generating the embedding vector includes:

generating the embedding vector using embedding function including an embedding layer and a plurality of fully-connected layers.

13 . The method of claim 10 , wherein

the denoising model includes a cascaded structure of residual blocks (RBs) and convolutional layers and uses the cascaded structure to generate the filtered image area, and

each RB is a convolutional block having a skip path between an input and an output.

14 . The method of claim 13 , wherein generating the filtered image area comprises:

multiplying a feature generated by a preset convolutional layer among the convolutional layers by an absolute value of the embedding vector.

15 . A method of storing a bitstream of a video into a non-transitory computer-readable recording medium, wherein the bitstream is generated by a video encoding method, and the video encoding method comprises:

obtaining an image area having been reconstructed and a quantization parameter of the image area;

generating an embedding vector based on the quantization parameter and a prediction type of the image area; and

generating a filtered image area based on the embedding vector by using a denoising model that is based on deep learning,

wherein the prediction type of the image area indicates one of an intra prediction type in which the image area is predicted independently, a predictive type in which the image area is predicted based on a reference image area in a single direction, or a bi-predictive type in which the image area is predicted based on at least one reference image area in bi-directions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: KANG, JE WON; KIM, NA YOUNG; LEE, JUNG KYUNG; PARK, SEUNG WOOK
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION; EWHA UNIVERSITY - INDUSTRY COLLABORATION FOUNDATION
Reel/Frame 065186/0640 →
Priority Claims (2)
KR 10-2021-0042090 · Mar 31, 2021 · national
KR 10-2022-0036249 · Mar 23, 2022 · national
Continuity (2)
Continuation PCTKR2022004171 · Mar 24, 2022
Related Publication 20240031580A1 · Jan 25, 2024
References Cited (28)
US 9288485B2 · Chono et al. · 2016 [cited by applicant]
US 9445102B2 · Schwaab et al. · 2016 [cited by applicant]
US 9601125B2 · Atti et al. · 2017 [cited by applicant]
US 9899032B2 · Atti et al. · 2018 [cited by applicant]
US 11057627B2 · Kim et al. · 2021 [cited by applicant]
US 11553186B2 · Kim et al. · 2023 [cited by applicant]
US 20130010859A1 · Schwaab et al. · 2013 [cited by applicant]
US 20130121408A1 · Chono et al. · 2013 [cited by applicant]
US 20130208808A1 · Sasai · 2013 [cited by examiner]
US 20140071233A1 · Lim · 2014 [cited by examiner]
US 20140229172A1 · Atti et al. · 2014 [cited by applicant]
US 20170148460A1 · Atti et al. · 2017 [cited by applicant]
US 20190095795A1 · Ren · 2019 [cited by examiner]
US 20200029080A1 · Kim et al. · 2020 [cited by applicant]
US 20200374547A1 · Gao · 2020 [cited by examiner]
US 20210250597A1 · Du · 2021 [cited by examiner]
US 20210297675A1 · Kim et al. · 2021 [cited by applicant]
US 20220148130A1 · Tang · 2022 [cited by examiner]
US 20230106301A1 · Kim et al. · 2023 [cited by applicant]
CN 110062234A · 2019 [cited by examiner]
KR 20130054318A · 2013 [cited by applicant]
KR 20190123288A · 2019 [cited by applicant]
KR 20190127090A · 2019 [cited by applicant]
KR 20210034103A · 2021 [cited by applicant]
International Search Report and Written Opinion cited in corresponding PCT application No. PCT/KR2022/004171; Jul. 4, 2022; 11 pp. [cited by applicant]
Ren Yang et al., Multi-Frame Quality Enhancement for Compressed Video, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv:1803.04680; 2018, 10 pp. [cited by applicant]
Xu, Xiangyu, Muchen Li, and Wenxiu Sun; Learning deformable kernels for image and video denoising; arXiv preprint arXiv:1904.06903; 2019, 10 pp. [cited by applicant]
Yang, Fuzhi, et al.; Learning texture transformer network for image super-resolution; arXiv:2006.04139; 2020, 22 pp. [cited by applicant]