IP Library › Granted Patent US 12,238,343
Granted Patent B2
US 12,238,343 · App. 18/193,763 · Granted Feb 25, 2025

Video coding with neural network based in-loop filtering

Inventors: Tsung-Chuan Ma (Beijing, CN); Wei Chen (Beijing, CN); Xiaoyu Xiu (Beijing, CN); Yi-Wen Chen (Beijing, CN); Hong-Jheng Jhu (Beijing, CN); Che-Wei Kuo (Beijing, CN); Xianglin Wang (Bejing, CN); Bing Yu (Beijing, CN)
Assignee: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
H04N19/82H04N19/186
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,343
App. No.
18/193,763
Granted
Feb 25, 2025
Kind
B2
Abstract

An electronic apparatus performs a method of decoding video data, including: reconstructing, from a video bitstream, a picture frame that includes a luma component, a first and a second chroma components, and applying a trained neural network based in-loop filter to the reconstructed picture frame by: converting a first resolution of the samples of the at least one of the first and the second chroma components to a second resolution of the samples of the luma component when the first resolution is different from the second resolution; concatenating samples of at least one of the first and the second chroma components with the luma component; processing the concatenated samples using a convolutional neural network; and reconverting the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution.

Claims (51)

1. A method of decoding video data, comprising:

reconstructing, from a video bitstream, a picture frame that includes a luma component, a first chroma component, and a second chroma component, and

applying a trained neural network based in-loop filter to the reconstructed picture frame by performing operations comprising:

concatenating samples of at least one of the first and the second chroma components with the luma component to create concatenated samples; and

processing the concatenated samples using a convolutional neural network

wherein the trained neural network based in-loop filter has been trained to obtain model parameters by reducing difference between output reconstructed picture frames from the convolutional neural network and corresponding ground-truth picture frames,

wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises:

converting a first resolution of the samples of the at least one of the first and the second chroma components to a second resolution of the samples of the luma component when the first resolution of the at least one of the first and the second chroma components is different from the second resolution of the luma component; and

reconverting the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution.

2. The method according to claim 1 , wherein the converting the first resolution of the samples of the at least one of the first and the second chroma components to a second resolution of the samples of the luma component, comprises:

up-sampling the samples of at least one of the first and the second chroma components from the first resolution to the second resolution of the samples of the luma component, and

wherein the reconverting the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution comprises:

down-sampling the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution.

3. The method according to claim 1 , wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises:

feeding quantization parameter (QP) information of prediction residuals as an additional input to the convolutional neural network.

4. The method according to claim 3 , wherein the QP information is normalized QP value, or

the QP information is QP step size, or

the QP information is QP maps, each of the QP maps combined with a corresponding color plane of the luma component, the first chroma component, and the second chroma component, respectively.

5. The method according to claim 1 , wherein the corresponding ground-truth picture frames are available at an encoder.

6. The method according to claim 5 , wherein the model parameters are explicitly signaled from the encoder to a decoder.

7. The method according to claim 1 , wherein the trained neural network based in-loop filter has been trained for each encoded picture frame and updated frequently.

8. The method according to claim 1 , wherein the trained neural network based in-loop filter has been trained for each encoded picture frame and updated frequently with overfitting enabled.

9. The method according to claim 1 , wherein the trained neural network based in-loop filter has been trained and saved offline for next use.

10. The method according to claim 1 , wherein the trained neural network based in-loop filter is-has been trained and saved offline for next use with overfitting disabled.

11. The method according to claim 1 , wherein the converting and the reconverting are processed by convolution modules.

12. The method according to claim 1 , wherein the trained neural network based in-loop filter has been trained from reconstructed picture frame training data with a deblocking filter applied.

13. The method according to claim 1 , wherein the trained neural network based in-loop filter has been trained from reconstructed picture frame training data without a deblocking filter applied.

14. The method according to claim 1 , wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises: segmenting the reconstructed picture frame into a plurality of overlapping blocks, wherein an overlapping size of the overlapping blocks is selected according to one or more of a content of the reconstructed picture frame, width to height ratio of the overlapping blocks, and width to height ratio of the reconstructed picture frame.

15. The method according to claim 1 , wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises: rounding sample values at an output of the neural network by applying an offset value, wherein the offset value is a fixed value or a value signaled and determined from an encoder.

16. The method according to claim 1 , wherein the convolutional neural network is ResNet.

17. An electronic apparatus comprising:

one or more processing units;

memory coupled to the one or more processing units; and

a plurality of programs stored in the memory that, when executed by the one or more processing units, cause the electronic apparatus to perform decoding video data by performing operations comprising:

reconstructing, from a video bitstream, a picture frame that includes a luma component, a first chroma component, and a second chroma component, and

applying a trained neural network based in-loop filter to the reconstructed picture frame by performing operations comprising:

concatenating samples of at least one of the first and the second chroma components with the luma component to create concatenated samples; and

processing the concatenated samples using a convolutional neural network,

wherein the trained neural network based in-loop filter has been trained to obtain model parameters by reducing difference between output reconstructed picture frames from the convolutional neural network and corresponding ground-truth picture frames,

wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises:

converting a first resolution of the samples of the at least one of the first and the second chroma components to a second resolution of the samples of the luma component when the first resolution of the at least one of the first and the second chroma components is different from the second resolution of the luma component; and

reconverting the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution.

18. A non-transitory computer readable storage medium storing a plurality of programs for execution by an electronic apparatus having one or more processing units, wherein the plurality of programs, when executed by the one or more processing units, cause the electronic apparatus to perform decoding video data comprising by performing operations comprising:

reconstructing, from a video bitstream, a picture frame that includes a luma component, a first chroma component, and a second chroma component, and

applying a trained neural network based in-loop filter to the reconstructed picture frame by performing operations comprising:

concatenating samples of at least one of the first and the second chroma components with the luma component to create concatenated samples; and

processing the concatenated samples using a convolutional neural network

wherein the trained neural network based in-loop filter has been trained to obtain model parameters by reducing difference between output reconstructed picture frames from the convolutional neural network and corresponding ground-truth picture frames,

wherein applying the trained neural network based in-loop filter to the reconstructed picture frame further comprises:

converting a first resolution of the samples of the at least one of the first and the second chroma components to a second resolution of the samples of the luma component when the first resolution of the at least one of the first and the second chroma components is different from the second resolution of the luma component; and

reconverting the samples of the at least one of the first and the second chroma components processed by the convolutional neural network from the second resolution back to the first resolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2024
From: MA, TSUNG-CHUAN; CHEN, WEI; XIU, XIAOYU; CHEN, YI-WEN; JHU, HONG-JHENG; KUO, CHE-WEI; WANG, XIANGLIN; YU, BING
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 068746/0839 →
Continuity (3)
Continuation PCTUS2021052915 · Sep 30, 2021
Provisional Application 63086538 · Oct 1, 2020
Related Publication 20230319314A1 · Oct 5, 2023
References Cited (14)
US 20190108618A1 · Hwang et al. · 2019 [cited by applicant]
US 20190246102A1 · Cho · 2019 [cited by examiner]
US 20190273948A1 · Yin · 2019 [cited by examiner]
US 20200213587A1 · Galpin · 2020 [cited by examiner]
US 20200252654A1 · Su et al. · 2020 [cited by applicant]
US 20210136416A1 · Kim · 2021 [cited by examiner]
EP 3930323A1 · 2021 [cited by examiner]
WO 2020180737A1 · 2020 [cited by applicant]
International Search Report and Written Opinion for PCT/US2021/052915 mailed Jan. 14, 2022, 6 pages. [cited by applicant]
Extended European Search Report corresponding to European Application No. 21876491.8 (11 pages) (Oct. 15, 2024). [cited by applicant]
Hsiao et al. “AHG9: Convolutional neural network loop filter” Joint Video Exports Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 (Jan. 2019). [cited by applicant]
Hsiao et al. “CE10-1. 2: Convolutional neural network loop filter” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 (Jul. 2019). [cited by applicant]
Huang et al. “Frame-Wise CNN-Based Filtering for Intra-Frame Quality Enhancement of HEVC Videos” IEEE Transactions on Circuits and Systems for Video Technology 31(6):2100-2113 (Jun. 2021). [cited by applicant]
Wang et al. “Attention-Based Dual-Scale CNN In-Loop Filter for Versatile Video Coding” IEEE Access 7:145214-145226 (Sep. 30, 2019). [cited by applicant]
Cited By (3)
US 12,555,201 US 12,610,047 US 12,744,915