IP Library Granted Patent US 12,659,474
Granted Patent B2
US 12,659,474 · App. 18/487,350 · Granted Jun 16, 2026

Unified neural network filter model

Inventors: Yue Li (San Diego, CA); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Junru Li (Beijing, CN); Meng Wang (Beijing, CN); Siwei Ma (Beijing, CN); Shiqi Wang (Kowloon, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.; BYTEDANCE (HK) LIMITED
H04N19/117G06T9/002H04N19/186H04N19/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,659,474
App. No.
18/487,350
Granted
Jun 16, 2026
Kind
B2
Abstract

A method implemented by a video coding apparatus. The method includes applying a first filter to an unfiltered sample of a video unit to generate a filtered sample. The first filter is a neural network (NN) filter based on a non-deep learning-based filter (NDLF) being disabled, and the first filter is the NDLF based on the NN filter being disabled. The method also includes performing a conversion between a video media file and a bitstream based on the filtered sample that was generated.

Claims (37)

1 . A method for processing video data, comprising:

determining, for a conversion between a video unit of a video and a bitstream of the video, interaction between a first filtering tool and a second filtering tool for being applied to the video unit, wherein the first filtering tool comprises a neural network (NN) filter, and the second filtering tool comprises a non-deep learning-based filtering (NDLF) tool; and

performing the conversion based on the determining,

wherein the interaction between the first filtering tool and the second filtering tool depends on a color format or a color component of the video unit, or

wherein the interaction between the first filtering tool and the second filtering tool depends on a profile flag, a tier flag, a level flag, or a constraint flag.

2 . The method of claim 1 , wherein when the NN filter is applied to the video unit, one or more filters of the NDLF tool are disabled for the video unit.

3 . The method of claim 2 , wherein the NDLF tool comprises a deblocking (DB) filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), a cross-component (CC)-SAO filter, a CC-ALF, a luma mapping with chroma scaling (LMCS) filter, a bilateral filter, or a transform-domain filter.

4 . The method of claim 3 , wherein the NN filter is applied to the video unit based on the ALF being disabled.

5 . The method of claim 3 , wherein the NN filter is applied to a chroma component of the video unit based on the CC-ALF being disabled, or

wherein the NN filter is applied to the chroma component of the video unit based on the CC-SAO filter being disabled.

6 . The method of claim 2 , wherein information related to the one or more filters of the NDLF tool is not included in the bitstream.

7 . The method of claim 2 , wherein information related to the one or more filters of the NDLF tool is included in the bitstream, and

wherein the information comprises an indication indicating that the one or more filters of the NDLF tool are disabled for the video unit when the NN filter is applied to the video unit.

8 . The method of claim 1 , wherein when the NN filter is applied to the video unit and information related to one or more filters of the NDLF tool is not included in the bitstream, one or more filters of the NDLF tool are determined to be disabled for the video unit.

9 . The method of claim 1 , wherein when the NN filter is applied to the video unit, all types of filters of the NDLF tool are disabled for the video unit.

10 . The method of claim 1 , wherein the video unit comprises a coding tree unit (CTU), a coding tree block (CTB), a CTU row, a CTB row, a slice, a tile, a picture, a sequence, or a subpicture.

11 . The method of claim 1 , wherein when the NN filter is disabled for the video unit, one or more filters of the NDLF tool are applied to the video unit.

12 . The method of claim 1 , wherein when one or more filters of the NDFL tool are not applied to the video unit, the NN filter is applied to the video unit.

13 . The method of claim 1 , wherein both of the NN filter and the NDFL tool are applied to the video unit, and

wherein the NN filter is applied before or after one or more filters of the NDFL tool.

14 . The method of claim 1 , wherein the NN filter is applied to the video unit after an in-loop filtering tool.

15 . The method of claim 1 , wherein the NN filter is applied to the video unit before an in-loop filtering tool,

wherein whether the in-loop filtering tool is applied to the video is determined at a slice level, a picture level, a subpicture level, or a tile level, and

wherein the in-loop filtering tool comprises a sample adaptive offset (SAO) filter, a cross-component (CC)-SAO filter, or an adaptive loop filter (ALF).

16 . The method of claim 1 , wherein the conversion comprises encoding the video into the bitstream.

17 . The method of claim 1 , wherein the conversion comprises decoding the video from the bitstream.

18 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor cause the processor to:

determine, for a conversion between a video unit of a video and a bitstream of the video, interaction between a first filtering tool and a second filtering tool for being applied to the video unit, wherein the first filtering tool comprises a neural network (NN) filter, and the second filtering tool comprises a non-deep learning-based filtering (NDLF) tool; and

perform the conversion based on the determining,

wherein the interaction between the first filtering tool and the second filtering tool depends on a color format or a color component of the video unit, or

wherein the interaction between the first filtering tool and the second filtering tool depends on a profile flag, a tier flag, a level flag, or a constraint flag.

19 . A method for storing a bitstream of a video, comprising:

determining, for a video unit of the video, interaction between a first filtering tool and a second filtering tool for being applied to the video unit, wherein the first filtering tool comprises a neural network (NN) filter, and the second filtering tool comprises a non-deep learning-based filtering (NDLF) tool;

generating the bitstream based on the determining; and

storing the bitstream in a non-transitory computer-readable recording medium,

wherein the interaction between the first filtering tool and the second filtering tool depends on a color format or a color component of the video unit, or

wherein the interaction between the first filtering tool and the second filtering tool depends on a profile flag, a tier flag, a level flag, or a constraint flag.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: LI, YUE; ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 065408/0048 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: WANG, MENG; MA, SIWEI
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065408/0104 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: WANG, SHIQI
To: BYTEDANCE (HK) LIMITED
Reel/Frame 065408/0160 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: LI, JUNRU
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065408/0188 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065408/0206 →
Priority Claims (1)
WO PCT/CN2021/087228 · Apr 14, 2021 · international
Continuity (2)
Continuation PCTCN2022086899 · Apr 14, 2022
Related Publication 20240056570A1 · Feb 15, 2024
References Cited (50)
US 11218695B2 · Park · 2022 [cited by applicant]
US 11949918B2 · Li · 2024 [cited by applicant]
US 11979591B2 · Li et al. · 2024 [cited by applicant]
US 20200154145A1 · Du et al. · 2020 [cited by applicant]
US 20200327702A1 · Wang · 2020 [cited by applicant]
US 20200382793A1 · Gao · 2020 [cited by applicant]
US 20210021820A1 · Ikai et al. · 2021 [cited by applicant]
US 20210044811A1 · Hodgkinson · 2021 [cited by examiner]
US 20220038721A1 · Li · 2022 [cited by applicant]
US 20220101095A1 · Li · 2022 [cited by examiner]
US 20220103864A1 · Wang · 2022 [cited by applicant]
US 20220295116A1 · Ma · 2022 [cited by applicant]
CN 103891293A · 2014 [cited by applicant]
CN 107197260A · 2017 [cited by applicant]
CN 108184129A · 2018 [cited by applicant]
CN 110971915A · 2020 [cited by applicant]
CN 111052740A · 2020 [cited by applicant]
CN 111064958A · 2020 [cited by applicant]
CN 111133756A · 2020 [cited by applicant]
CN 111194555A · 2020 [cited by applicant]
CN 111541894A · 2020 [cited by applicant]
CN 111866522A · 2020 [cited by applicant]
CN 112422993A · 2021 [cited by applicant]
WO 2019182159A1 · 2019 [cited by applicant]
WO 2020177072A1 · 2020 [cited by applicant]
WO 2021051369A1 · 2021 [cited by applicant]
WO 2021211966A1 · 2021 [cited by applicant]
Bossen Ed., et al., “VTM Software Manual,” Joint Video Experts Team (JVET) of ITUT SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Document:: JVET Software Manual, Aug. 13, 2020, 46 Pages. [cited by applicant]
Document: JVET-S2001-vH, Bross, B., et al., “Versatile Video Coding (Draft 10),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 19th Meeting: by teleconference, Jun. 22-Jul. 1, 2020, 5… [cited by applicant]
Retrieved from the internet: https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/-/tags/VTM-10.0, Jan. 10, 2024, 1 page. [cited by applicant]
Document: JVET-L0147, Lim, S., et al., “CE2: Subsampled Laplacian calculation (Test 6.1, 6.2, 6.3, and 6.4),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WVG 11 12th Meeting: Macao, CN, O… [cited by applicant]
Document: JVET-N0242, Taquet, J., et al., “CE5: Results of tests CE5-3.1 to CE5-3.4 on Non-Linear Adaptive Loop Filter.,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 14th Meeting: G… [cited by applicant]
Balle, J., et al., “End-to-end optimization of nonlinear transform codes for perceptual quality,” In PCS. IEEE 2016, 5 pages. [cited by applicant]
Theis, L., et al., “Lossy image compression with compressive autoencoders,” Published as a conference paper at ICLR 2017, arXiv:1703.00395v1 [stat.ML], Mar. 1, 2017, 19 pages. [cited by applicant]
Li, J., et al., “Fully Connected Network-Based Intra Prediction for Image Coding,” IEEE Transactions on Image Processing, 2018, 11 pages. [cited by applicant]
Dai, Y., et al., “A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding,” In MMM. Springer, arXiv:1608.06690v2 [cs.MM], Oct. 29, 2016, 12 pages. [cited by applicant]
Song, R., et al., “Neural Network-Based Arithmetic Coding of Intra Prediction Modes in HEVC,” In VCIP. IEEE, Dec. 10-13, 2017, 4 pages. [cited by applicant]
Pfaff, J., et al., “Neural network based intra prediction for video coding,” In Applications of Digital Image Processing XLI, vol. 10752. International Society for Optics and Photonics, 1075213, Sep. 2018, 7 pages. [cited by applicant]
Document: JVET-U0068-v2, Li, Y., et al., “AHG11: Convolutional Neural Network-based In-Loop Filter with Adaptive Model Selection,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 21st Meeting… [cited by applicant]
Document: JVET-U2016-r1, Liu, S., et al., “JVET common test conditions and evaluation procedures for neural network-based video coding technology,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/S… [cited by applicant]
Timofte, R., et al., Retrieved from the internet: https://data.vision.ee.ethz.ch/cvl/DIV2K/ R, Jan. 10, 2024, 6 pages. [cited by applicant]
Ma, D., et al., “BVI-DVC: A Training Database for Deep Video Compression,” arXiv:2003.13552v2 [eess.IV], Oct. 8, 2020, 11 pages. [cited by applicant]
Retrieved from the internet: https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/-/tags/VTM-11.0, Jan. 10, 2024, 1 page. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/714,014 dated Jun. 23, 2023, 27 pages. [cited by applicant]
International Search Report from PCT Application No. PCT/CN2022/086899 dated Jul. 11, 2022, 8 pages. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/720,125 dated Jun. 8, 2023, 27 pages. [cited by applicant]
Office Action from U.S. Appl. No. 18/657,026 dated Jul. 30, 2025, 27 pages. [cited by applicant]
First Office Action for Chinese Application No. 202210357676.0, mailed Apr. 23, 2025, 25 pages. [cited by applicant]
First Office Action for Chinese Application No. 202210399919.7, mailed on Apr. 18, 2025, 20 pages. [cited by applicant]
Li Y., et al., “AHG11: Convolutional Neural Network-based In-Loop Filter with Adaptive Model Selection,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21st Meeting, by teleconference, Jan.… [cited by applicant]