IP Library Granted Patent US 12,363,350
Granted Patent B2
US 12,363,350 · App. 18/257,227 · Granted Jul 15, 2025

Neural network based in-loop filtering for video coding

Inventors: Zhao Wang (Beijing, CN); Changyue Ma (Beijing, CN); Ru-Ling Liao (Beijing, CN); Yan Ye (San Diego, CA)
Assignee: Alibaba Group Holding Limited
H04N19/86H04N19/176H04N19/59
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,363,350
App. No.
18/257,227
Granted
Jul 15, 2025
Kind
B2
Abstract

The present disclosure provides methods for performing training and executing of a multi-density neural network in video processing. An exemplary method comprises: receiving a video stream comprising a plurality of pictures; processing the plurality of pictures using a first branch of a first block in the neural network, wherein the neural network is configured to reduce blocking artifacts in video compression of the video stream and the first branch comprises one or more residual blocks; and processing the plurality of pictures using a second branch of the first block in the neural network, wherein the second branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks.

Claims (50)

1. A method for using a neural network in video processing, comprising:

receiving a video stream comprising a plurality of pictures;

processing the plurality of pictures using a first branch of a first block in the neural network, wherein the neural network is configured to reduce blocking artifacts in video compression of the video stream and the first branch comprises one or more residual blocks and the first branch is configured to maintain full resolution;

processing the plurality of pictures using a second branch of the first block in the neural network, wherein the second branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks; and

merging an output of the first branch of the first block and an output of the second branch of the first block.

2. The method according to claim 1 , wherein merging the output of the first branch of the first block and the output of the second branch of the first block further comprising:

performing an element-wise dot product on the output of the first branch and the output of the second branch.

3. The method according to claim 1 , wherein:

the down-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions; and

the up-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions including a sigmoid function.

4. The method according to claim 1 , further comprising:

processing the plurality of pictures using a third branch of the first block of the neural network, wherein the third branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks, and the down-sampling processing of the third branch has a sampling size that is different from a sampling size of the down-sampling processing of the second branch.

5. The method according to claim 1 , further comprising:

processing the plurality of pictures using a second block of the neural network, wherein the second block comprises a plurality of branches.

6. The method according to claim 1 , further comprising:

performing pre-processing on the plurality of pictures, wherein the pre-processing comprises performing a mean-shift subtraction on the plurality of pictures; and

performing post-processing on the plurality of pictures, wherein the post-processing comprises a mean-shift add on training pictures and the mean-shift add corresponds to the mean-shift subtraction in the pre-processing.

7. A system for using a neural network in video processing, the system comprising:

a memory storing a set of instructions; and

a processor configured to execute the set of instructions to cause the system to perform:

receiving a video stream comprising a plurality of pictures;

processing the plurality of pictures using a first branch of a first block in the neural network, wherein the neural network is configured to reduce blocking artifacts in video compression of the video stream, the first branch comprises one or more residual blocks and the first branch is configured to maintain full resolution;

processing the plurality of pictures using a second branch of the first block in the neural network, wherein the second branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks; and

merging an output of the first branch of the first block and an output of the second branch of the first block.

8. The system according to claim 7 , wherein the processor is further configured to execute the set of instructions to cause the system to perform:

performing an element-wise dot product on the output of the first branch and the output of the second branch.

9. The system according to claim 7 , wherein:

the down-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions; and

the up-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions including a sigmoid function.

10. The system according to claim 7 , wherein the processor is further configured to execute the set of instructions to cause the system to perform:

processing the plurality of pictures using a third branch of the first block of the neural network, wherein the third branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks, and the down-sampling processing of the third branch has a sampling size that is different from a sampling size of the down-sampling processing of the second branch.

11. The system according to claim 7 , wherein the processor is further configured to execute the set of instructions to cause the system to perform:

processing the plurality of pictures using a second block of the neural network, wherein the second block comprises a plurality of branches.

12. The system according to claim 7 , wherein the processor is further configured to execute the set of instructions to cause the system to perform:

performing pre-processing on the plurality of pictures, wherein the pre-processing comprises performing a mean-shift subtraction on the plurality of pictures; and

performing post-processing on the plurality of pictures, wherein the post-processing comprises a mean-shift add on training pictures and the mean-shift add corresponds to the mean-shift subtraction in the pre-processing.

13. A non-transitory computer readable medium that stores a set of instructions that is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing training of a neural network in video processing, the method comprising:

receiving a video stream comprising a plurality of pictures;

processing the plurality of pictures using a first branch of a first block in the neural network, wherein the neural network is configured to reduce blocking artifacts in video compression of the video stream, the first branch comprises one or more residual blocks and the first branch is configured to maintain full resolution;

processing the plurality of pictures using a second branch of the first block in the neural network, wherein the second branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks; and

merging an output of the first branch of the first block and an output of the second branch of the first block.

14. The non-transitory computer readable medium of claim 13 , wherein the method further comprises:

performing an element-wise dot product on the output of the first branch and the output of the second branch.

15. The non-transitory computer readable medium of claim 13 , wherein:

the down-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions; and

the up-sampling processing of the second branch comprises one or more convolutional layers including a 1×1 convolutional layer, and one or more activation functions including a sigmoid function.

16. The non-transitory computer readable medium of claim 13 , wherein the method further comprises:

processing the plurality of pictures using a third branch of the first block of the neural network, wherein the third branch comprises a down-sampling processing, an up-sampling processing, and one or more residual blocks, and the down-sampling processing of the third branch has a sampling size that is different from a sampling size of the down-sampling processing of the second branch.

17. The non-transitory computer readable medium of claim 13 , wherein the method further comprises:

processing the plurality of pictures using a second block of the neural network, wherein the second block comprises a plurality of branches.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2026
From: ALIBABA INNOVATION PRIVATE LIMITED
To: SIM IP 5 LLC
Reel/Frame 075529/0713 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2026
From: ALIBABA GROUP HOLDING LIMITED
To: ALIBABA INNOVATION PRIVATE LIMITED
Reel/Frame 075522/0643 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: WANG, ZHAO; MA, CHANGYUE; LIAO, RULING; YE, YAN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 064433/0349 →
Continuity (1)
Related Publication 20240048777A1 · Feb 8, 2024
References Cited (35)
US 20120148103A1 · Hampel · 2012 [cited by examiner]
US 20150195573A1 · Aflaki Beni et al. · 2015 [cited by applicant]
US 20170286780A1 · Zhang · 2017 [cited by examiner]
US 20210012537A1 · Xu · 2021 [cited by examiner]
US 20220094962A1 · Choi · 2022 [cited by examiner]
US 20220148131A1 · Wang · 2022 [cited by examiner]
US 20220191553A1 · Auyeung · 2022 [cited by examiner]
CN 102196272A · 2011 [cited by applicant]
CN 110120019A · 2019 [cited by applicant]
CN 110188776A · 2019 [cited by applicant]
CN 111164651A · 2020 [cited by applicant]
“JVET software repository,” https://https://jvet.hhi.fraunhofer.de/svn/syn_HMJEMSoftware/.:https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/branches/HM-13.0-QTBT/. [cited by applicant]
Balle et al., End-to-End Optimized Image Compression, ICLR, 2017, 27 pages. [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 7),”JVET-P2001-vE, 16th Meeting: Geneva, CH, Oct. 1-11, 2019, 488 pages. [cited by applicant]
Dai et al., “A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding,” in Proc. MMM, Reykjavik, Iceland, Jan. 2017, pp. 28-39. [cited by applicant]
Fu et al., “Sample Adaptive Offset in the HEVC Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, p. 1755-1764 (2012). [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778) (2019). [cited by applicant]
He et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1026-1034. [cited by applicant]
International Telecommunications Union “Series H: Audiovisual and Multimedia Systems Infrastructure of audiovisual services—Coding of moving video”, ITU-T Telecommunication Standardization Sector of ITU, Apr. 2013, 317 … [cited by applicant]
Jia et al., “Spatial-Temporal Residue Network Based In-Loop Filter for Video Coding,” 2017 IEEE Visual Communications and Image Processing (VCIP), St. Petersburg, FL, 2017, pp. 1-4. [cited by applicant]
Kang et al., “Multi-model/multi-scale convolutional neural network based in-loop filter design for next generation video codec,” in Proc. IEEE ICIP, 2017, pp. 26-30. [cited by applicant]
Li et al., “A deep learning approach for multi-frame in-loop filter of HEVC.” IEEE Transactions on Image Processing, vol. 28, No. 11, Nov. 2019, pp. 5663-5678. [cited by applicant]
Lu et al. “DVC: An end-to-end deep video compression framework,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 11006-11015, 2019. [cited by applicant]
Meng et al., “A new HEVC in-loop filter based on multi-channel long-short-term dependency residual networks,” in Proc. DCC, 2018, pp. 187-196. [cited by applicant]
Nair et al., “Rectified Linear Units Improve Restricted Boltzmann Machines,” in Proceedings of the 27th international conference on machine learning, CML-10, pp. 807-814, 2010. [cited by applicant]
Norkin et al., “HEVC Deblocking Filter, IEEE Transactions on Circuits and Systems for Video Technology,” vol. 22, No. 12, pp. 1746-1754 (2012). [cited by applicant]
Park et al., “CNN-based in-loop filtering for coding efficiency improvement,” Proc. IEEE IVMSP Workshop, Jul. 2016, pp. 1-5. [cited by applicant]
PCT International Search Report and Written Opinion mailed Oct. 27, 2021, issued in corresponding International Application No. PCT/CN2021/072774 (7 pgs.). [cited by applicant]
Svoboda et al., “Compression Artifacts Removal Using Convolutional Neural Networks,” Journal of WSCG, vol. 24, No. 2, pp. 65-70, 2016. [cited by applicant]
Wang et al., “Attention-Based Dual-Scale CNN In-Loop Filter for Versatile Video Coding,” IEEE Access, vol. 7, 145214-145226, 2019. [cited by applicant]
Zhang et al., “Residual non-local attention networks for image restoration,” arXiv preprint arXiv:1903.10082, 2019. [cited by applicant]
European Patent Office Communication issued for Application No. 21920203.3 the Supplementary European Search Report (Art. 153(7) EPC) and the European search opinion dated Sep. 2, 2024, 32 pages. [cited by applicant]
Wang et al., “AHG11: Multi-density network for in-loop filtering,” JVET-U0055, 21 [cited by applicant]
Alshina et al., EE Summary Report: Neural Network-based Video Coding, JVET-U0023_r1, 21 [cited by applicant]
Wang et al., “EE-1.6: Neural network based in-loop filtering,” JVET-U0054, 21 [cited by applicant]