IP Library › Granted Patent US 12,549,776
Granted Patent B2
US 12,549,776 · App. 18/488,782 · Granted Feb 10, 2026

Using neural network filtering in video coding

Inventors: Yue Li (San Diego, CA); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA)
H04N19/85G06N3/04H04N19/107H04N19/124H04N19/174H04N19/176H04N19/184H04N19/573H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,549,776
App. No.
18/488,782
Granted
Feb 10, 2026
Kind
B2
Abstract

Methods, systems, apparatus for media processing are described. One example method of digital media processing includes determining, for a conversion between visual media data and a bitstream of the visual media data, how to apply one or more convolutional neural network filters to at least some samples of a video unit of the visual media data according to a rule; and performing the conversion based on the determining.

Claims (35)

1 . A method of processing visual media data, comprising:

determining, for a conversion between visual media data and a bitstream of the visual media data, how to apply one or more convolutional neural network filters to samples of a video unit of the visual media data according to a rule; and

performing the conversion based on the determining,

wherein when the one or more convolutional neural network filters are applied, a difference between a convolutional neural network filtered sample and its unfiltered version is clipped to a range,

wherein the rule specifies that a selection of a set of convolutional neural network filters depends on a group of pictures (GOP) size of the video unit, and

wherein the video unit is a slice or a picture or a tile or a subpicture or a coding tree block or a coding tree unit.

2 . The method of claim 1 , wherein a convolutional neural network filer is implemented using a convolutional neural network.

3 . The method of claim 1 , wherein the rule specifies that the determining is based on decoded information associated with a video unit of the visual media data, wherein the decoded information includes at least one of prediction modes, transform types, a skip flag, or coded block flag (CBF) values.

4 . The method of claim 1 , wherein the rule specifies that information related to the one or more convolutional neural network filters is controlled at a granularity smaller than a video unit of the visual media data.

5 . The method of claim 4 , wherein the information is controlled at a sample or pixel level.

6 . The method of claim 4 , wherein the information is controlled at a sample row, a sample column, or a sample line level.

7 . The method of claim 4 , wherein the rule specifies that a set of convolutional neural network filters is determined based on a value of a sample or a position of the sample within the video unit of the visual media data.

8 . The method of claim 1 , wherein the rule specifies that a selection of a set of convolutional neural network filters depends on a temporal layer identification of a video unit of the visual media data.

9 . The method of claim 8 , wherein the temporal layer identification is classified to more than one category and a given set of convolutional neural network filters is applied for a corresponding category.

10 . The method of claim 9 , wherein classification of the temporal layer identification is based on the GOP size.

11 . The method of claim 1 , wherein the rule specifies that a set of convolutional neural network filters is utilized for video units with different temporal layers.

12 . The method of claim 1 , wherein the performing of the conversion comprises generating the bitstream from the visual media data.

13 . The method of claim 1 , wherein the performing of the conversion comprises generating the visual media data from the bitstream.

14 . An apparatus for processing visual media data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine, for a conversion between visual media data and a bitstream of the visual media data, how to apply one or more convolutional neural network filters to samples of a video unit of the visual media data according to a rule; and

perform the conversion based on the determining,

wherein when the one or more convolutional neural network filters are applied, a difference between a convolutional neural network filtered sample and its unfiltered version is clipped to a range,

wherein the rule specifies that a selection of a set of convolutional neural network filters depends on a group of pictures (GOP) size of the video unit, and

wherein the video unit is a slice or a picture or a tile or a subpicture or a coding tree block or a coding tree unit.

15 . The apparatus of claim 14 , wherein the rule specifies that the determining is based on decoded information associated with a video unit of the visual media data, wherein the decoded information includes at least one of prediction modes, transform types, a skip flag, or coded block flag (CBF) values.

16 . The apparatus of claim 14 , wherein the rule specifies that information related to the one or more convolutional neural network filters is controlled at a granularity smaller than a video unit of the visual media data.

17 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine, for a conversion between visual media data and a bitstream of the visual media data, how to apply one or more convolutional neural network filters to samples of a video unit of the visual media data according to a rule; and

perform the conversion based on the determining,

wherein when the one or more convolutional neural network filters are applied, a difference between a convolutional neural network filtered sample and its unfiltered version is clipped to a range,

wherein the rule specifies that a selection of a set of convolutional neural network filters depends on a group of pictures (GOP) size of the video unit, and

wherein the video unit is a slice or a picture or a tile or a subpicture or a coding tree block or a coding tree unit.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the rule specifies that the determining is based on decoded information associated with a video unit of the visual media data, wherein the decoded information includes at least one of prediction modes, transform types, a skip flag, or coded block flag (CBF) values.

19 . The non-transitory computer-readable storage medium of claim 17 , wherein the rule specifies that information related to the one or more convolutional neural network filters is controlled at a granularity smaller than a video unit of the visual media data.

20 . The apparatus of claim 14 , wherein the rule specifies that a selection of a set of convolutional neural network filters depends on a temporal layer identification of a video unit of the visual media data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: LI, YUE; ZHANG, LI; ZHANG, KAI
To: BYTEDANCE INC.
Reel/Frame 065408/0437 →
Continuity (3)
Continuation 17488179 · Sep 28, 2021
Provisional Application 63087113 · Oct 2, 2020
Related Publication 20240048775A1 · Feb 8, 2024
References Cited (117)
US 6674911B1 · Pearlman · 2004 [cited by examiner]
US 10341670B1 · Brailovskiy · 2019 [cited by examiner]
US 10701394B1 · Caballero et al. · 2020 [cited by applicant]
US 10706350B1 · Tran et al. · 2020 [cited by applicant]
US 11023730B1 · Zhou et al. · 2021 [cited by applicant]
US 11412225B2 · Lee · 2022 [cited by examiner]
US 12022098B2 · Li · 2024 [cited by applicant]
US 20070206680A1 · Bourge · 2007 [cited by examiner]
US 20130182950A1 · Morales · 2013 [cited by applicant]
US 20140192862A1 · Flynn · 2014 [cited by applicant]
US 20150317557A1 · Julian et al. · 2015 [cited by applicant]
US 20160014411A1 · Sychev · 2016 [cited by applicant]
US 20170213321A1 · Matviychuk · 2017 [cited by applicant]
US 20170330363A1 · Song et al. · 2017 [cited by applicant]
US 20170332098A1 · Rusanovskyy · 2017 [cited by applicant]
US 20170347060A1 · Wang et al. · 2017 [cited by applicant]
US 20180204051A1 · Li et al. · 2018 [cited by applicant]
US 20180352264A1 · Guo · 2018 [cited by examiner]
US 20180359491A1 · Kang · 2018 [cited by applicant]
US 20190035116A1 · Xing et al. · 2019 [cited by applicant]
US 20190108432A1 · Lu et al. · 2019 [cited by applicant]
US 20190147372A1 · Luo et al. · 2019 [cited by applicant]
US 20190171935A1 · Agrawal · 2019 [cited by applicant]
US 20190246102A1 · Cho et al. · 2019 [cited by applicant]
US 20190273948A1 · Yin et al. · 2019 [cited by applicant]
US 20200057935A1 · Wang et al. · 2020 [cited by applicant]
US 20200267392A1 · Lu · 2020 [cited by applicant]
US 20200311870A1 · Jung · 2020 [cited by applicant]
US 20210044834A1 · Li · 2021 [cited by examiner]
US 20210064985A1 · Sun · 2021 [cited by applicant]
US 20210084318A1 · Kuo · 2021 [cited by examiner]
US 20210125380A1 · Lee · 2021 [cited by applicant]
US 20210158072A1 · Zhang et al. · 2021 [cited by applicant]
US 20210195223A1 · Chang · 2021 [cited by applicant]
US 20210375260A1 · Yu et al. · 2021 [cited by applicant]
US 20210385443A1 · Masule · 2021 [cited by examiner]
US 20210400311A1 · Hsiao · 2021 [cited by examiner]
US 20220046236A1 · Li · 2022 [cited by examiner]
US 20220094919A1 · Lai · 2022 [cited by examiner]
US 20220101095A1 · Li et al. · 2022 [cited by applicant]
US 20220103816A1 · Karczewicz · 2022 [cited by applicant]
US 20220109890A1 · Li et al. · 2022 [cited by applicant]
US 20220116600A1 · Rosewarne et al. · 2022 [cited by applicant]
US 20220141458A1 · Sakurai · 2022 [cited by examiner]
US 20220201328A1 · Galpin · 2022 [cited by examiner]
US 20220256145A1 · Huang · 2022 [cited by examiner]
US 20220295116A1 · Ma · 2022 [cited by examiner]
US 20220303587A1 · Lai · 2022 [cited by examiner]
US 20220312006A1 · Taquet · 2022 [cited by examiner]
US 20220337848A1 · Francois · 2022 [cited by applicant]
CN 105144719A · 2015 [cited by applicant]
CN 106911930A · 2017 [cited by applicant]
CN 107197260A · 2017 [cited by applicant]
CN 107483930A · 2017 [cited by applicant]
CN 107888927A · 2018 [cited by applicant]
CN 107925762A · 2018 [cited by applicant]
CN 108184129A · 2018 [cited by applicant]
CN 108353182A · 2018 [cited by applicant]
CN 109565594A · 2019 [cited by applicant]
CN 110062226A · 2019 [cited by applicant]
CN 110263841A · 2019 [cited by applicant]
CN 110506277A · 2019 [cited by applicant]
CN 110740319A · 2020 [cited by applicant]
CN 110971915A · 2020 [cited by applicant]
CN 111405283A · 2020 [cited by applicant]
CN 111757122A · 2020 [cited by applicant]
CN 114339221B · 2024 [cited by applicant]
EP 3342164A1 · 2018 [cited by applicant]
EP 3451293A1 · 2019 [cited by applicant]
JP 2007281634A · 2007 [cited by applicant]
WO 2016199330A1 · 2016 [cited by applicant]
WO 2017036370A1 · 2017 [cited by applicant]
WO 2019182159A1 · 2019 [cited by applicant]
WO 2020062074A1 · 2020 [cited by applicant]
WO 202057787A1 · 2020 [cited by applicant]
Retrieved from the internet: http://phenix.it-sudparis.eu/jvet/doc_end_user/current_document.phpid=10399, Jan. 12, 2024, 1 page. [cited by applicant]
Retrieved from the internet: https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/-/tags/VTM-10.0, Jan. 12, 2024, 1 page. [cited by applicant]
Lim, S., et al., “CE2: Subsampled Laplacian calculation (Test 6.1, 6.2, 6.3, and 6.4),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, JVET-L… [cited by applicant]
Balle, J., et al., “End-to-end optimization of nonlinear transform codes for perceptual quality,” In PCS. IEEE, 2016, 5 pages. [cited by applicant]
Theis, L., et al., “Lossy image compression with compressive autoencoders,” Published as a conference paper at ICLR 2017, arXiv:1703.00395v1 [stat.ML], Mar. 1, 2017, 19 pages. [cited by applicant]
Li, J., et al., “Fully Connected Network-Based Intra Prediction for Image Coding,” IEEE Transactions on Image Processing, 2018, 11 pages. [cited by applicant]
Dai, Y., et al., “A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding,” In MMM. Springer, arXiv:1608.06690v2 [cs.MM], Oct. 29, 2016, 12 pages. [cited by applicant]
Song, R., et al., “Neural Network-Based Arithmetic Coding of Intra Prediction Modes in HEVC,” In VCIP. IEEE, Dec. 10-13, 2017, 4 pages. [cited by applicant]
Pfaff, J., et al., “Neural network based intra prediction for video coding,” In Applications of Digital Image Processing XLI, vol. 10752. International Society for Optics and Photonics, 1075213, Sep. 2018, 7 pages. [cited by applicant]
“CNN-Based In-Loop Filtered for Coding Efficiency Improvement,” IEEE 12th Image, Video, and Multidimensional Signal Processing Workshop, XP032934608, Jul. 11, 2016, 5 pages. [cited by applicant]
Xu, L., et al., “Non-CE10: A CNN based in-loop filter for intra frame,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 15th Meeting: Gothenburg, SE, Jul. 3-12, 2019, JVET-00157, 5 pag… [cited by applicant]
Bross, B., et al., “Versatile Video Coding (Draft 5),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH, Mar. 19-27, 2019, JVET-N1001-v10, 406 pages. [cited by applicant]
Li, Y., “AHG11: Convolutional neural networks-based in-loop filter,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 20th Meeting, by teleconference, Oct. 7-16, 2020, JVET-T0088, 4 pag… [cited by applicant]
Hsiao, Y., “CE13-1.1: Convolutional neural network loop filter,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH, Mar. 19-27, 2019, JVET-N0110-v1, 5 pages. [cited by applicant]
Haiyan, W, et al., “Research on video coding loop filtering technology based on progressive network” Jun. 2020, 8 pages. [cited by applicant]
Taquet, J., “Results of tests CE5-3.1 to CE5-3.4 on Non-Linear Adaptive Loop Filter,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, CH, Mar. 19-27, 2019, JVET-N… [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/488,179 dated Aug. 18, 2022, 20 pages. [cited by applicant]
Notice of Allowance from U.S. Appl. No. 17/488,179 dated Jun. 8, 2023, 8 pages. [cited by applicant]
Final Office Action from U.S. Appl. No. 17/488,179 dated Dec. 22, 2022, 10 pages. [cited by applicant]
Extended European Search Report from European Application No. 21199983.4 dated Mar. 4, 2022, 11 pages. [cited by applicant]
Extended European Search Report from European Application No. 21199996.6 dated Mar. 3, 2022, 11 pages. [cited by applicant]
Wang, Y., “Research on Convolutional Neural Network based In-Loop Filtering Technique for Video Compression,” May 2019, 67 pages. [cited by applicant]
Chen, T., et al., “DeepCoder: A deep neural network based video compression,” 2017 IEEE Visual Communications and Image Processing (VCIP), Dec. 10-13, 2017, 4 pages. [cited by applicant]
Chinese Office Action from U.S. Patent Application 202111163049.5 dated Mar. 19, 2024, 7 pages. [cited by applicant]
Advanced Video Coding for Generic Audiovisual Services, Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video, ITU-H.264, Jun. 2019, 836 pages. [cited by applicant]
Audio Visual and Multimedia Systems Infrastructure of Audio Visual Services—Coding of Moving Video—High Efficiency Video Coding, Series H, ITU-T Reccommendation H.265, Feb. 2018, 692 pages. [cited by applicant]
Transmission of Non-Telephone Signals, Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Video, ITU-T H.262, Jul. 1995, 211 pages. [cited by applicant]
Video Coding for Low Bit Rate Communication, Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video, ITU-T H.263, Jan. 2005, 226 pages. [cited by applicant]
Bossen, ED., et al., “VTM Software Manual,” Joint Video Experts Team (JVET) of ITUT SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Document: JVET Software Manual, Aug. 13, 2020, 46 pages. [cited by applicant]
Bossen, et al., “JVET common test conditions and software reference configurations for SDR video,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-L1010-v1, 12th Meeting… [cited by applicant]
JVET-S2001-vH-Brose, et al., “Versatile Video Coding (Draft 10),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document, 19th Meeting: by teleconference, Jun. 22-Jul. 1, 2020, 548 p… [cited by applicant]
Liu, et al., “Methodology and reporting template for neural network coding tool test,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, Document: JVET-0041-v4, 20th Meeting: by teleconference… [cited by applicant]
Ma, et al., “BVI-DVC: A Training Database for Deep Video Compression,” ArXiv:2003.13552, Oct. 8, 2020, 11 pages. [cited by applicant]
Timofte, et al., DIV2K dataset: DIVerse 2K Resolution High Quality Images as Used for the Challenges @NTIRE (CVPR 2017 (https://www.vision.ee.ethz.ch/ntire17) and CPVR 2018 (http://www.vision.ee.ethz.ch/ntire18)) and @P… [cited by applicant]
“Series H: Audiovisual And Multimedia Systems Infrastrcuture of Audiovisual Services—Coding of Moving Video,” Telecommunication Standardization Sector of ITU, Infrastructure of Audiovisual Services—Coding of Moving Vide… [cited by applicant]
JVET-S2001-v1-Bross, et al., “Versatlie Video Coding (Draft 10),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document, 19th Meeting: by teleconference, Jun. 22-Jul. 1, 2020, 548 p… [cited by applicant]
Yan-Dong, L., et al., “Deblocking Filter in Video Coding,” Communications Technology, vol. 41, No. 201, Sep. 2008, 6 pages. With English Translation. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 202210210562.3 dated Jul. 29, 2025, 15 pages. [cited by applicant]
Document: JVET-O0157-v2, Xu, L., et al., “Non-CE10: A CNN based in-loop filter for intra frame,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 15th Meeting: Gothenburg, SE, Jul. 3-12,… [cited by applicant]
Document: JVET-N0110-v2, Hsiao, Y., et al., “CE13-1.1: Convolutional neural network loop filter,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 14th Meeting: Geneva, CH, Mar. 19-27, 2… [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 2021111630442 dated Nov. 23, 2025, 11 pages. [cited by applicant]
European Office Action from European Patent Application No. 21199983.4 dated Dec. 4, 2025, 7 pages. [cited by applicant]