IP Library › Granted Patent US 12,621,499
Granted Patent B2
US 12,621,499 · App. 18/624,889 · Granted May 5, 2026

Unified neural network in-loop filter signaling

Inventors: Yue Li (San Diego, CA); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Junru Li (Beijing, CN); Meng Wang (Beijing, CN); Siwei Ma (Beijing, CN); Shiqi Wang (Hong Kong, CN)
Assignees: Lemon Inc.; Beijing Bytedance Network Technology Co., Ltd.; Bytedance Inc.; Bytedance (HK) Limited
H04N19/82H04N19/117H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,499
App. No.
18/624,889
Granted
May 5, 2026
Kind
B2
Abstract

A method implemented by a video coding apparatus includes applying a neural network (NN) filter to an unfiltered sample of a video unit to generate a filtered sample. The NN filter is applied based on a syntax element of the video unit. The method also includes converting between a video media file and a bitstream based on the filtered sample that was generated.

Claims (37)

1 . A method implemented by a video coding apparatus, comprising:

applying a neural network (NN) filter to an unfiltered sample of a video unit to generate a filtered sample, wherein the NN filter is applied based on a syntax element of the video unit; and

converting between a video media file and a bitstream based on the filtered sample that was generated,

wherein the syntax element indicates at least one selected from the group consisting of: whether to enable the NN filter, a number of NN filters to be applied, and a type of NN filter to be applied,

wherein a first level comprises a sequence level, and a second level comprises a picture level, and

wherein the syntax element is a first syntax element at the first level that indicates whether a NN filter can be adaptively selected at the second level to be applied to a picture or a slice of the video unit.

2 . The method of claim 1 , wherein:

a syntax element indicated in the first level is indicated in a sequence parameter set (SPS) and/or a sequence header of the video unit;

a syntax element indicated in the second level is indicated in a picture header, a picture parameter set (PPS), and/or a slice header of the video unit; and

a third level comprises a subpicture level a syntax element indicated in the third level is indicated for a patch of the video unit, a coding tree unit (CTU) of the video unit, a coding tree block (CTB) of the video unit, a block of the video unit, a subpicture of the video unit, a tile of the video unit, a slice of the video unit, or a region of the video unit.

3 . The method of claim 2 , wherein the syntax element is the first syntax element further at the second level that is conditionally applied based on a second syntax element at the first level, wherein the NN filter is applied at the second level based on the first syntax element based on the second syntax element being a flag that is true, and wherein the NN filter is not applied based on the second syntax element being false.

4 . The method of claim 2 , wherein the syntax element is the first syntax element further at the second level that indicates whether a NN filter can be adaptively selected at the third level to be applied to a subpicture of the video unit, or that indicates whether usage of the NN filter can be controlled at the third level.

5 . The method of claim 2 , wherein the syntax element is the first syntax element further at the second level that indicates whether a NN filter can be adaptively selected at the second level, used at the second level, or applied on the second level; and wherein the first syntax element is signaled based on an indication that the NN filter can be adaptively selected at the second level or an indication that a number of NN filters is greater than one.

6 . The method of claim 2 , wherein the syntax element is the first syntax element further at the third level that is conditionally applied based on a second syntax element at the first level and/or a third syntax at the second level, wherein the first syntax element is coded using context, wherein the NN filter is applied at the third level based on the first syntax element based on one of the second syntax element and the third syntax element being a flag that is true, and wherein the NN filter is not applied based on one of the second syntax element and the third syntax element being false.

7 . The method of claim 2 , wherein the syntax element is the first syntax element further at the third level that is signaled based on an indication that the NN filter can be adaptively selected at the third level or an indication that a number of NN filters is greater than one.

8 . The method of claim 2 , wherein the syntax element is signaled responsive to the NN filter being enabled for a picture or a slice of the video unit, and wherein the NN filter is one of a plurality (T) of NN filters, and wherein the syntax element includes an index (k).

9 . The method of claim 8 , further comprising applying the k th NN filter at the second level of the video unit based on the index k>=0 and k<T.

10 . The method of claim 8 , further comprising adaptively selecting a NN filter at the third level based on the index k>=T.

11 . The method of claim 8 , wherein the index k is restricted to be in a range from 0 to (T−1).

12 . The method of claim 2 , wherein the syntax element is coded based on a context model that is selected based on a number of allowed NN filters, wherein a filter model index for a color component of the video unit is configured to specify one of K context models, and wherein the one of the K context models is specified as a minimum of K−1 and binIdx, wherein binIdx is an index of a bin to be coded.

13 . The method of claim 12 , wherein a filter model index for first and second color components of the video unit is coded with a same set of contexts.

14 . The method of claim 12 , wherein the filter model index for a first color component of the video unit is coded with a different set of contexts than the filter model index for a second color component of the video unit.

15 . The method of claim 1 , wherein the syntax element is signaled using context coding or bypass coding, or is binarized using fixed-length coding, unary coding, truncated unary coding, signaled unary coding, signed truncated unary coding, truncated binary coding, or exponential Golomb coding.

16 . The method of claim 1 , wherein the conversion comprises generating the bitstream according to the video media file.

17 . The method of claim 1 , wherein the conversion comprises parsing the bitstream to obtain the video media file.

18 . An apparatus for coding video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor cause the processor to:

apply a neural network (NN) filter to an unfiltered sample of a video unit to generate a filtered sample, wherein the NN filter is applied based on a syntax element of the video unit; and

convert between a video media file and a bitstream based on the filtered sample that was generated,

wherein the syntax element indicates at least one selected from the group consisting of: whether to enable the NN filter, a number of NN filters to be applied, and a type of NN filter to be applied,

wherein a first level comprises a sequence level, and a second level comprises a picture level, and

wherein the syntax element is a first syntax element at the first level that indicates whether a NN filter can be adaptively selected at the second level to be applied to a picture or a slice of the video unit.

19 . A non-transitory computer readable medium storing a bitstream of a video that is generated by a method performed by a video processing apparatus, wherein the method comprises:

applying a neural network (NN) filter to an unfiltered sample of a video unit to generate a filtered sample, wherein the NN filter is applied based on a syntax element of the video unit; and

generating the bitstream based on the filtered sample that was generated,

wherein the syntax element indicates at least one selected from the group consisting of: whether to enable the NN filter, a number of NN filters to be applied, and a type of NN filter to be applied,

wherein a first level comprises a sequence level, and a second level comprises a picture level, and

wherein the syntax element is a first syntax element at the first level that indicates whether a NN filter can be adaptively selected at the second level to be applied to a picture or a slice of the video unit.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: WANG, MENG; MA, SIWEI
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 067446/0121 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: LI, JUNRU
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 067448/0332 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: WANG, SHIQI
To: BYTEDANCE (HK) LIMITED
Reel/Frame 067448/0464 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: LI, YUE; ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 067449/0080 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 067449/0175 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2024
From: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.; BYTEDANCE (HK) LIMITED
To: LEMON INC.; BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.; BYTEDANCE (HK) LIMITED
Reel/Frame 067449/0298 →
Continuity (3)
Continuation 17720125 · Apr 13, 2022
Provisional Application 63176871 · Apr 19, 2021
Related Publication 20240276020A1 · Aug 15, 2024
References Cited (50)
US 11218695B2 · Park · 2022 [cited by applicant]
US 11949918B2 · Li · 2024 [cited by applicant]
US 11979591B2 · Li et al. · 2024 [cited by applicant]
US 20200154145A1 · Du · 2020 [cited by applicant]
US 20200327702A1 · Wang · 2020 [cited by applicant]
US 20200382793A1 · Gao · 2020 [cited by applicant]
US 20210021820A1 · Ikai · 2021 [cited by applicant]
US 20210044811A1 · Hodgkinson · 2021 [cited by applicant]
US 20220038721A1 · Li · 2022 [cited by applicant]
US 20220101095A1 · Li · 2022 [cited by applicant]
US 20220103864A1 · Wang · 2022 [cited by applicant]
US 20220295116A1 · Ma · 2022 [cited by applicant]
CN 103891293A · 2014 [cited by applicant]
CN 107197260A · 2017 [cited by applicant]
CN 108184129A · 2018 [cited by applicant]
CN 110971915A · 2020 [cited by applicant]
CN 111052740A · 2020 [cited by applicant]
CN 111064958A · 2020 [cited by applicant]
CN 111133756A · 2020 [cited by applicant]
CN 111194555A · 2020 [cited by applicant]
CN 111541894A · 2020 [cited by applicant]
CN 111866522A · 2020 [cited by applicant]
CN 112422993A · 2021 [cited by applicant]
WO 2019182159A1 · 2019 [cited by applicant]
WO 2020177072A1 · 2020 [cited by applicant]
WO 2021051369A1 · 2021 [cited by applicant]
WO WO2021211966A1 · 2021 [cited by examiner]
Bross, et al., “Versatile Video Coding (Draft 10),” Document: JVET-S2001-vH, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 19th Meeting: by teleconference, Jun. 22-Jul. 1, 2020, 548 … [cited by applicant]
Bossen, Ed., et al., “VTM Software Manual,” Document: JVET-Software Manual, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Aug. 13, 2020, 46 pages. [cited by applicant]
Lim, et al, “CE2: Subsampled Laplacian calculation (Test 6.1, 6.2, 6.3, and 6.4),” Document JVET-L0147, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-… [cited by applicant]
Taquet, et al., “CE5: Results of tests CE5-3.1 to CE5-3.4 on Non-Linear Adaptive Loop Filter,” Document JVET-N0242, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva… [cited by applicant]
Balle, et. al., “End-to-end optimization of nonlinear transform codes for perceptual quality,” In PCTS. IEEE, 2016, pp. 1-5. [cited by applicant]
Theis, et. al., “Lossy Image Compression with Compressive Autoencoders,” ICLR 2017, arXiv:1703.00395v1 [stat.MIL], Mar. 1, 2017, 19 pages. [cited by applicant]
Li, et. al., “Fully Connected Network-Based Intra Prediction for Image Coding,” IEEE Transaction on Image Processing, Vo., 27, Issue 7, 2018, pp. 3236-3247. [cited by applicant]
Dai, et. al., “A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding,” MMM, Springer, arXiv:1608.06690v2 [cs.MM], Oct. 29, 2016,, pp. 28-39. [cited by applicant]
Song, et. al., “Neural Network-Based Arithmetic Coding of Intra Prediction Modes in HEVC,” VCIP 2017, IEEE, Dec. 10-13, 2017, 4 pages. [cited by applicant]
Pfaff, et al., “Neural network based intra prediction for video coding,” In Applications of Digital Image Processing XLI, vol. 10752. International Society for Optics and Photonics, 1075213, 7 pages. [cited by applicant]
Li, et al, “AHG11: Convolutional Neural Network-based In-Loop Filter with Adaptive Model Selection,” Document: JVET-U0068-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21st Meeting, by… [cited by applicant]
Liu, et al., “JVET common test conditions and evaluation procedures for neural network-based video coding technology,” Document: JVET-U2016-r1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29… [cited by applicant]
Timofte, et al., “DIV2K dataset: DIVerse 2K resolution high quality images as used for the challenges @ NTIRE (CVPR 2017 and CVPR 2018) and @ PIRM (ECCV 2018),” https://data.vision.ee.ethz.ch/cvl/DIV2K/, 6 pages. [cited by applicant]
Ma, et al., “BVI-DVC: A Training Database for Deep Video Compression,” arXiv:2003.13552v2 [33ss.IV] Oct. 8, 2020, 11 pages. [cited by applicant]
Suehring, K., Retrieved from the internet: https://vcgit.hhi .fraunhofer.de/jvet/VVCSoftware_VTM/-tags/VTM-10.0, Jul. 15, 2022, 2 pages. [cited by applicant]
Bossen, F., Retrieved from the internet: https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/-tags/VTM-11.0, Jul. 15, 2022, 1 page. [cited by applicant]
Foreign Communication From A Related Counterpart Application, International Application No. PCT/CN2022/086899, International Search Report dated Jul. 11, 2022, 8 pages. [cited by applicant]
Non-Final Office Action dated Jun. 23, 2023, 27 pages, U.S. Appl. No. 17/714,014, filed Jun. 23, 2023. [cited by applicant]
Non-Final Office Action dated Jun. 8, 2023, 27 pages, U.S. Appl. No. 17/720,125, filed Apr. 13, 2022. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 202210357676.0 dated Apr. 23, 2025, 25 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 202210399919.7 dated Apr. 18, 2025, 20 pages. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 18/487,350 dated Jun. 18, 2025, 30 pages. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 18/657,026 dated Jul. 30, 2025, 27 pages. [cited by applicant]