IP Library › Granted Patent US 12,323,608
Granted Patent B2
US 12,323,608 · App. 17/714,027 · Granted Jun 3, 2025

On neural network-based filtering for imaging/video coding

Inventors: Yue Li (San Diego, CA); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA)
Assignee: Lemon Inc
H04N19/436H04N19/117H04N19/124H04N19/132H04N19/136H04N19/146H04N19/1883H04N19/31H04N19/70H04N19/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,608
App. No.
17/714,027
Granted
Jun 3, 2025
Kind
B2
Abstract

A method of processing video data. The method includes selecting an in-loop filter from a plurality of neural network (NN) filter model candidates, wherein the plurality of NN filter model candidates are based on a reconstructed quality level of a video unit, and performing a conversion between a video media file comprising the video unit and a bitstream based on the in-loop filter selected. A corresponding video coding apparatus and non-transitory computer readable medium are also disclosed.

Claims (35)

1. A method of processing video data, comprising:

selecting an in-loop filter from a plurality of neural network (NN) filter model candidates, wherein the plurality of NN filter model candidates are based on a reconstructed quality level of a video unit; and

performing a conversion between a video media file comprising the video unit and a bitstream based on the in-loop filter selected,

wherein the method further comprises determining whether a group of NN filter model candidates are the same or different for video units across different temporal layers,

wherein syntax elements corresponding to the in-loop filter selected are coded in the bitstream before syntax elements corresponding to an adaptive loop filter (ALF), and

wherein the method further comprises determining that a same index in a bitstream is associated with different NN filters for two video units.

2. The method of claim 1 , wherein the one or more NN filter model candidates comprise one or more pretrained convolutional neural network (CNN) filter models.

3. The method of claim 1 , wherein each of the plurality of NN filter model candidates corresponds to a different reconstructed quality level of video unit.

4. The method of claim 1 , wherein the reconstructed quality level of the video unit corresponds to a quantization parameter (QP) of the video unit or at least one of a constant rate factor and a bitrate of the video unit.

5. The method of claim 1 , wherein:

the in-loop filter selected is one of a plurality of in-loop filters including a second in-loop filter, and wherein application of the second in-loop filter is dependent on whether or how the in-loop filter selected is applied; and

the plurality of in-loop filters comprise at least one of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a cross-component adaptive loop filter (CCALF), and a bilateral filter.

6. The method of claim 1 , wherein syntax elements corresponding to the in-loop filter selected are coded in the bitstream at a coding tree unit (CTU) level before syntax elements corresponding to an adaptive loop filter (ALF) or before syntax elements corresponding to a sample adaptive offset (SAO) filter.

7. The method of claim 1 , wherein the in-loop filter selected is coded in a supplemental enhancement information (SEI) message of the bitstream.

8. The method of claim 1 , wherein the same index is disposed in a supplemental enhancement information (SEI) message in the bitstream.

9. The method of claim 1 , wherein the determination of whether the group of NN filter model candidates are the same or different for video units across different temporal layers is specified in a rule included in a supplemental enhancement information (SEI) message of the bitstream.

10. The method of claim 1 , wherein a rule included in the bitstream specifies that a first subgroup of the group of NN filter model candidates is to be used in an in-loop filtering operation across a first subgroup of the different temporal layers, and that a second subgroup of the group of NN filter model candidates is to be used in an in-loop filtering operation across a second subgroup of the different temporal layers.

11. The method of claim 10 , wherein the first subgroup of the different temporal layers comprises layers having a temporal index of no greater than K1, and wherein at least one of the one or more NN filter model candidates to be used in the in-loop filtering operation across the first subgroup is specified by a rule included in the bitstream based on a number of intra coded samples of the first subgroup.

12. The method of claim 10 , wherein a rule included in the bitstream associates the group of NN filter model candidates with both a first temporal layer and a separate second temporal layer of the different temporal layers.

13. The method of claim 10 , wherein a rule included in the bitstream associates the group of NN filter model candidates with a specific temporal layer of the different temporal layers.

14. The method of claim 1 , wherein the conversion includes encoding the video media file into the bitstream.

15. The method of claim 1 , wherein the conversion includes decoding the video media file from the bitstream.

16. An apparatus for coding video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor cause the processor to:

select an in-loop filter from a plurality of neural network (NN) filter model candidates, wherein the plurality of NN filter model candidates are based on a reconstructed quality level of a video unit; and

convert between a video media file comprising the video unit and a bitstream based on the in-loop filter selected,

wherein the instructions, upon execution by the processor, further cause the processor to determine whether a group of NN filter model candidates are the same or different for video units across different temporal layers,

wherein syntax elements corresponding to the in-loop filter selected are coded in the bitstream before syntax elements corresponding to an adaptive loop filter (ALF), and

wherein the processor is further caused to determine that a same index in a bitstream is associated with different neural network (NN) filters for two video units.

17. A method for storing a bitstream of a video, comprising:

selecting an in-loop filter from a plurality of neural network (NN) filter model candidates, wherein the plurality of NN filter model candidates are based on a reconstructed quality level of a video unit;

performing a conversion between a video media file comprising the video unit and a bitstream based on the in-loop filter selected; and

storing the bitstream in a non-transitory computer-readable recording medium,

wherein the method further comprises determining whether a group of NN filter model candidates are the same or different for video units across different temporal layers,

wherein syntax elements corresponding to the in-loop filter selected are coded in the bitstream before syntax elements corresponding to an adaptive loop filter (ALF), and

wherein the method further comprises determining that a same index in a bitstream is associated with different neural network (NN) filters for two video units.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2023
From: LI, YUE; ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 062566/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 062622/0516 →
Continuity (2)
Provisional Application 63171743 · Apr 7, 2021
Related Publication 20220337853A1 · Oct 20, 2022
References Cited (22)
US 20180249158A1 · Huang · 2018 [cited by examiner]
US 20200244997A1 · Galpin · 2020 [cited by examiner]
US 20200304836A1 · Li · 2020 [cited by examiner]
US 20220103864A1 · Wang · 2022 [cited by examiner]
US 20220217403A1 · Choi · 2022 [cited by examiner]
US 20220256227A1 · Rezazadegan Tavakoli · 2022 [cited by examiner]
US 20220295116A1 · Ma · 2022 [cited by examiner]
US 20220321919A1 · Deshpande · 2022 [cited by examiner]
US 20220337857A1 · Choi · 2022 [cited by examiner]
US 20230328293A1 · Chen · 2023 [cited by examiner]
US 20230345003A1 · Chen · 2023 [cited by examiner]
Li et al. “AHG11: Convolutional Neural Network-based In-Loop Filter with Adaptive Model Selection”. Jan. 5, 2021. (Year: 2021). [cited by examiner]
Bross, B., et al., “Versatile Video Coding (Draft 10),” http://phenix.it-sudparis.eu/jvet/doc_end_user/current_document.php?id=10399, Jul. 15, 2022, 1 page. [cited by applicant]
Suehring, K., https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/-/tags/VTM-10.0, Jul. 15, 2022, 2 pages. [cited by applicant]
Document: JVET-L0147, Lim, S-C., et al., “CE2: Subsampled Laplacian calculation (Test 6.1, 6.2, 6.3, and 6.4), ” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting: Macao, CN,… [cited by applicant]
Document: JVET-N0242, Taquet, J., et al., “E5: Results of tests CE5-3.1 to CE5-3.4 on Non-Linear Adaptive Loop Filter,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 14th Meeting: Gen… [cited by applicant]
Balle, J., et al., “End-to-end optimization of nonlinear transform codes for perceptual quality,” In PCS. IEEE, 2016, 5 pages. [cited by applicant]
Theis, L., et al., “Lossy image compression with compressive autoencoders,” Published as a conference paper at ICLR, arXiv preprint arXiv:1703.00395, Mar. 1, 2017, 19 pages. [cited by applicant]
Li, J., et al., “Fully Connected Network-Based Intra Prediction for Image Coding,” IEEE Transactions on Image Processing 27, 2018, pp. 3236-3247. [cited by applicant]
Dai, Y., “A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding,” arXiv:1608.06690v2, [cs.MM] Oct. 29, 2016, 12 pages. [cited by applicant]
Song, R., et al., “Neural Network-Based Arithmetic Coding of Intra Prediction Modes in HEVC,” In VCIP, IEEE, 2017, 4 pages. [cited by applicant]
Pfaff, J., et al., “Neural network based intra prediction for video coding,” In Applications of Digital Image Processing KLI, vol. 10752, International Society for Optics and Photonics, 1075213, Sep. 17, 2018, 7 pages. [cited by applicant]
Cited By (1)
US 12,744,915