IP Library › Granted Patent US 12,744,915
Granted Patent B2
US 12,744,915 · App. 18/795,899 · Granted Sep 22, 2026

Method and apparatus for video coding using an in-loop filter based on a transformer

Inventors: Je Won Kang (Seoul, KR); Jin Heo (Yongin-si, KR); Seung Wook Park (Yongin-si, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION; EWHA UNIVERSITY—INDUSTRY COLLABORATION FOUNDATION
H04N19/176H04N19/119H04N19/154H04N19/172H04N19/42H04N19/60H04N19/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,915
App. No.
18/795,899
Granted
Sep 22, 2026
Kind
B2
Abstract

A method and an apparatus are disclosed for video using an in-loop filter based on Transformer. The video coding method and the apparatus apply a current video block to an attention module of a Transformer, which is a deep learning model. The video coding method and the apparatus utilize the resultant Transformer-based in-loop filter.

Claims (55)

1 . A method performed by a video decoding device for enhancing a picture quality of a reconstructed frame, the method comprising:

obtaining an input region of a preset size from the reconstructed frame which is a reconstruction of an original frame and has been reconstructed in advance by the video decoding device; and

generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,

wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,

wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and

wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.

2 . The method of claim 1 , wherein:

the in-loop filter further comprises a first convolutional neural network (CNN); and

generating the enhanced video region further includes

generating an input feature by feeding the input region into the first CNN, and

converting, by using the K consecutive Transformer blocks, the input feature into the final output feature based on the attention operation.

3 . The method of claim 2 , wherein the in-loop filter further comprises a second CNN, and wherein generating the enhanced video region includes feeding the final output feature into the second CNN to generate the enhanced video region.

4 . The method of claim 2 , wherein converting the input feature includes:

partitioning an input feature of each of transformer blocks into patches, and applying an attenuation operation to each of the patches to generate an output feature for the input feature of each of the transformer block.

5 . The method of claim 4 , wherein converting the input feature includes:

setting each of the patches as a query;

calculating, based on a similarity between two patches used in the attention operation, a self-attention score for each of the patches, and attention scores between each of the patches and other patches; and

weighted summing the attention scores to generate an attention value for each of the patches.

6 . The method of claim 4 , wherein converting the input feature includes:

when two patches used in the attention operation are not equal in size, applying a padding to equalize the two patches in size.

7 . The method of claim 5 , wherein converting the input feature includes:

not calculating the attention scores when the two patches used in the attention operation are not present together in the input feature.

8 . The method of claim 5 , wherein converting the input feature includes:

not calculating the attention scores when the two patches used in the attention operation reside in different data processing units, which are coding tree units (CTUs) or virtual pipeline data units (VPDUs).

9 . The method of claim 4 , wherein the input feature includes:

not applying the attenuation operation to each of the patches when each of the patches falls outside a boundary of a preset data processing unit; and

not applying the attenuation operation to a partial region of each of the patches based on a mask that indicates the partial region outside the boundary when the partial region falls outside the boundary.

10 . The method of claim 4 , wherein converting the input feature includes:

not applying the attention operation to each of the patches when each of the patches is present in a later order of decoding than a current block; and

not applying the attenuation operation to a partial region of each of the patches based on a mask that indicates the partial region in the later order when the partial region is present in the later order.

11 . The method of claim 4 , wherein each of the Transformer blocks includes:

consecutive Transformer layers; and

one optional convolution layer,

wherein each of the consecutive Transformer layers includes one encoder layer and one decoder layer.

12 . The method of claim 11 , wherein each of the Transformer blocks is formed as a residual block of the consecutive Transformer layers by using a skip connection.

13 . A method performed by a video encoding device for enhancing a picture quality of a reconstructed frame, the method comprising:

obtaining an input region of a preset size from the reconstructed frame, which is a reconstruction of an original frame and has been reconstructed in advance by the video encoding device; and

generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,

wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,

wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and

wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.

14 . The method of claim 13 , wherein:

the in-loop filter further comprises a first convolutional neural network (CNN); and

generating the enhanced video region further includes

generating an input feature by feeding the input region into the first CNN, and

converting, by using the K consecutive Transformer blocks, the input feature into the final output feature based on the attention operation.

15 . The method of claim 14 , wherein:

the in-loop filter further comprises a second CNN; and

generating the enhanced video region includes feeding the final output feature into the second CNN to generate the enhanced video region.

16 . A non-transitory computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising:

obtaining an input region of a preset size from a reconstructed frame, which is a reconstruction of an original frame and has been reconstructed in advance by a video encoding device; and

generating an enhanced video region that approximates the original frame by feeding the input region into an in-loop filter that is deep learning-based,

wherein the in-loop filter comprises K consecutive Transformer blocks, and K is a natural number,

wherein generating the enhanced video region includes converting, by using the K consecutive Transformer blocks, an input image into a final output feature based on an attention operation, and

wherein at least one of the K consecutive Transformer blocks performs the attention operation based on a query vector, a key vector, and a value vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2024
From: KANG, JE WON; HEO, JIN; PARK, SEUNG WOOK
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION; EWHA UNIVERSITY – INDUSTRY COLLABORATION FOUNDATION
Reel/Frame 068214/0337 →
Priority Claims (2)
KR 10-2022-0022404 · Feb 21, 2022 · national
KR 10-2023-0007352 · Jan 18, 2023 · national
Continuity (2)
Continuation PCTKR2023001180 · Jan 26, 2023
Related Publication 20240397057A1 · Nov 28, 2024
References Cited (25)
US 11297341B2 · Erfurt · 2022 [cited by applicant]
US 11575885B2 · Kang · 2023 [cited by applicant]
US 11671591B2 · Zhang · 2023 [cited by applicant]
US 11936853B2 · Kang · 2024 [cited by applicant]
US 12238343B2 · Ma · 2025 [cited by examiner]
US 12321870B2 · Zou · 2025 [cited by examiner]
US 12323607B2 · Cricrì · 2025 [cited by examiner]
US 12323608B2 · Li · 2025 [cited by examiner]
US 12363350B2 · Wang · 2025 [cited by examiner]
US 20200029071A1 · Kang · 2020 [cited by applicant]
US 20200366918A1 · Erfurt · 2020 [cited by applicant]
US 20210400286A1 · Kale · 2021 [cited by applicant]
US 20220103864A1 · Wang · 2022 [cited by examiner]
US 20220217403A1 · Choi · 2022 [cited by examiner]
US 20220264087A1 · Zhang · 2022 [cited by applicant]
US 20230188708A1 · Kang · 2023 [cited by applicant]
US 20230291894A1 · Zhang · 2023 [cited by applicant]
US 20230345003A1 · Chen · 2023 [cited by examiner]
US 20240155110A1 · Kang · 2024 [cited by applicant]
KR 20220012393A · 2022 [cited by applicant]
WO 2021088951A1 · 2021 [cited by applicant]
Convolutional neural network based In-loop filter with adaptive model selection; Jan. 2021 (Year: 2021). [cited by examiner]
Deep In-loop filter with adaptive model selection; Yue Li; 2021; (Year: 2021). [cited by examiner]
International Search Report and Written Opinion cited in corresponding international patent application No. PCT/KR2023/001180; May 17, 2023; 10 pp. [cited by applicant]
Yue Li et al.; AHG11: Deep In-Loop Filter with Adaptive Model Selection and External Attention; Document No. JVET-W0100; Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29; Jul. 2021; 6 pp. [cited by applicant]