IP Library › Granted Patent US 12,327,384
Granted Patent B2
US 12,327,384 · App. 17/566,282 · Granted Jun 10, 2025

Multiple neural network models for filtering during video coding

Inventors: Hongtao Wang (San Diego, CA); Venkata Meher Satchit Anand Kotra (Munich, DE); Jianle Chen (San Diego, CA); Marta Karczewicz (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T9/002H04N19/11H04N19/619H04N19/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,384
App. No.
17/566,282
Granted
Jun 10, 2025
Kind
B2
Abstract

An example device for filtering decoded video data includes one or more processors configured to execute a neural network filtering unit to: receive data from one or more other units of the device, the data from the one or more other units of the device being different than data for a decoded picture of video data, and wherein to receive the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to receive boundary strength data from a deblocking unit of the device; determine one or more neural network models to be used to filter a portion of the decoded picture; and filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, including the boundary strength data.

Claims (79)

1. A method of filtering decoded video data, the method comprising:

receiving, by a neural network filtering unit of a video decoding device, data for a decoded picture of video data;

receiving, by the neural network filtering unit, data from one or more other units of the video decoding device, the data from the one or more other units being different than the data for the decoded picture, and wherein receiving the data from the one or more other units of the video decoding device comprises receiving boundary strength data from a deblocking unit of the video decoding device;

determining, by the neural network filtering unit, one or more neural network models to be used to filter a portion of the decoded picture; and

filtering, by the neural network filtering unit, the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device, including the boundary strength data.

2. The method of claim 1 , wherein receiving the data from the one or more other units of the video decoding device further comprises receiving the data from one or more of:

an intra-prediction unit of the video decoding device;

an inter-prediction unit of the video decoding device;

a transform processing unit of the video decoding device;

a quantization unit of the video decoding device;

a loop filter unit of the video decoding device;

a pre-processing unit of the video decoding device; or

a second neural network filtering unit of the video decoding device.

3. The method of claim 2 , wherein the loop filter unit comprises at least one of a sample adaptive offset (SAO) filtering unit or an adaptive loop filtering (ALF) unit.

4. The method of claim 1 , wherein receiving the data further comprises receiving one or more of coding unit (CU) partitioning data, prediction unit (PU) partitioning data, transform unit (TU) partitioning data, deblocking filtering data, quantization parameter (QP) data, intra-prediction data, inter-prediction data, data representing distance between the decoded picture and one or more reference pictures, or motion information for one or more decoded blocks of the decoded picture.

5. The method of claim 4 , wherein the deblocking filtering data includes one or more of whether long or short filters were used for deblocking or whether strong or weak filters were used for deblocking.

6. The method of claim 4 , wherein the intra-prediction data includes an intra-prediction mode.

7. The method of claim 4 , wherein the data representing the distance comprises data representing a difference between a picture order count (POC) value for the decoded picture and a POC value for a reference picture used to predict a block of the decoded picture.

8. The method of claim 1 , wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises providing the data from the one or more other units of the video decoding device as one or more additional input planes to a convolutional neural network (CNN).

9. The method of claim 1 , wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises:

combining a plurality of input planes into a combined input plane, including, for each position (i, j) of the plurality of input planes, setting a value for position (i, j) of the combined input plane equal to a maximum of the values at position (i, j) of the plurality of input planes; and

providing the combined input plane to a convolutional neural network (CNN).

10. The method of claim 1 , wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises adjusting output of the one or more neural network models using the data from the one or more other units of the video decoding device.

11. The method of claim 1 , further comprising adjusting the data from the one or more other units of the video decoding device prior to filtering the portion of the decoded picture.

12. The method of claim 11 , wherein adjusting the data comprises converting values of the data between integer representation and floating point representation.

13. The method of claim 11 , wherein adjusting the data comprises scaling values of the data to be within a range of values suitable for the one or more neural network models.

14. The method of claim 1 , wherein receiving the data comprises receiving partition data for the decoded picture, and wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises:

setting values at positions in an input plane collocated with positions of boundary samples defining partition boundaries in the decoded picture, as indicated by the partition data, to a first value;

setting values at positions in the input plane collocated with positions of internal samples that are non-boundary samples to a second value; and

filtering the portion of the decoded picture using the input plane as an input to at least one of the one or more neural network models.

15. The method of claim 14 , wherein the first value comprises 1 and the second value comprises 0.

16. The method of claim 14 , wherein the partition data comprises coding unit (CU) partition data and the input plane comprises a first partition plane, the method further comprising:

receiving prediction unit (PU) partition data;

forming a second input plane using the PU partition data;

receiving transform unit (TU) partition data; and

forming a third input plane using the TU partition data,

wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises filtering the portion of the decoded picture using the first partition plane, the second input plane, and the third input plane as inputs to at least one of the one or more neural network models.

17. The method of claim 1 , wherein filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device comprises:

converting deblocking filter data for the decoded picture from the deblocking unit to one or more input planes for at least one of the one or more neural network models; and

filtering the portion of the decoded picture using the one or more input planes as inputs to the at least one of the one or more neural network models.

18. The method of claim 1 , further comprising:

encoding a current picture; and

decoding the current picture to form the decoded picture.

19. The method of claim 18 , wherein determining the one or more neural network models comprises determining the one or more neural network models according to a rate-distortion computation.

20. The method of claim 1 , wherein the boundary strength data indicates that a boundary strength value is zero.

21. The method of claim 1 , wherein the boundary strength data indicates that a boundary strength value is one or two.

22. A device for filtering decoded video data, the device comprising:

a memory configured to store a decoded picture of video data; and

one or more processors implemented in circuitry and configured to execute a neural network filtering unit to:

receive data from one or more other units of the device, the data from the one or more other units of the device being different than data for the decoded picture, and wherein to receive the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to receive boundary strength data from a deblocking unit of the device;

determine one or more neural network models to be used to filter a portion of the decoded picture; and

filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, including the boundary strength data.

23. The device of claim 22 , wherein to receive the data from the one or more other units of the device, the one or more processors are further configured to execute the neural network filtering unit to receive data from one or more of:

an intra-prediction unit of the device;

an inter-prediction unit of the device;

a transform processing unit of the device;

a quantization unit of the device;

a loop filter unit of the device;

a pre-processing unit of the device; or

a second neural network filtering unit of the device.

24. The device of claim 22 , wherein to receive the data from the one or more other units of the device, the one or more processors are further configured to execute the neural network filtering unit to receive one or more of coding unit (CU) partitioning data, prediction unit (PU) partitioning data, transform unit (TU) partitioning data, deblocking filtering data, quantization parameter (QP) data, intra-prediction data, inter-prediction data, data representing distance between the decoded picture and one or more reference pictures, or motion information for one or more decoded blocks of the decoded picture.

25. The device of claim 22 , wherein to filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to provide the data from the one or more other units of the device as one or more additional input planes to a convolutional neural network (CNN).

26. The device of claim 22 , wherein to filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to adjust output of the one or more neural network models using the data from the one or more other units of the device.

27. The device of claim 22 , wherein to filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to adjust the data from the one or more other units of the device prior to filtering the portion of the decoded picture.

28. The device of claim 22 , wherein to filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the device, the one or more processors are configured to execute the neural network filtering unit to:

convert deblocking filter data for the decoded picture from the deblocking unit to one or more input planes for at least one of the one or more neural network models; and

filter the portion of the decoded picture using the one or more input planes as inputs to the at least one of the one or more neural network models.

29. The device of claim 22 , further comprising a display configured to display the decoded picture of the video data.

30. The device of claim 22 , wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

31. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause a processor of a video decoding device to execute a neural network filtering unit to:

receive data for a decoded picture of video data;

receive data from one or more other units of the video decoding device, the data from the one or more other units of the video decoding device being different than the data for the decoded picture, and wherein the instructions that cause the processor to receive the data from the one or more other units of the video decoding device comprise instructions that cause the processor to receive boundary strength data from a deblocking unit of the video decoding device;

determine one or more neural network models to be used to filter a portion of the decoded picture; and

filter the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device, including the boundary strength data.

32. A device for filtering decoded video data, the device comprising a filtering unit comprising:

means for receiving data for a decoded picture of video data;

means for receiving data from one or more other units of the video decoding device, the data from the one or more other units being different than the data for the decoded picture, and wherein the means for receiving the data from the one or more other units of the video decoding device comprises means for receiving boundary strength data from a deblocking unit of the video decoding device;

means for determining one or more neural network models to be used to filter a portion of the decoded picture; and

means for filtering the portion of the decoded picture using the one or more neural network models and the data from the one or more other units of the video decoding device, including the boundary strength data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2022
From: WANG, HONGTAO; KOTRA, VENKATA MEHER SATCHIT ANAND; CHEN, JIANLE; KARCZEWICZ, MARTA
To: QUALCOMM INCORPORATED
Reel/Frame 059067/0436 →
Continuity (2)
Provisional Application 63133733 · Jan 4, 2021
Related Publication 20220215593A1 · Jul 7, 2022
References Cited (27)
US 20190313095A1 · Ikeda · 2019 [cited by examiner]
US 20200213587A1 · Galpin et al. · 2020 [cited by applicant]
US 20200404278A1 · Ye et al. · 2020 [cited by applicant]
US 20200404335A1 · Egilmez et al. · 2020 [cited by applicant]
US 20210051320A1 · Tourapis et al. · 2021 [cited by applicant]
US 20220101095A1 · Li · 2022 [cited by examiner]
WO 2021051369A1 · 2021 [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/011021—ISA/EPO—Apr. 19, 2022, 14 pp. [cited by applicant]
Jia C., et al., “Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video Coding”, IEEE Transactions on Image Processing, IEEE, USA, vol. 28, No. 7, Jul. 1, 2019 (Jul. 1, 2019), pp. 3343… [cited by applicant]
Kawamura (KDDI) K., et al., “CE13-Related: Adaptive CNN Based in-Loop Filtering with Boundary Weights”, 14. JVET Meeting, Mar. 19, 2019-Mar. 27, 2019, Geneva, (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11 … [cited by applicant]
Li (Bytedance) Y., et al., “AHG11: Convolutional Neural Network-Based In-Loop Filter with Adaptive Model Selection”, 21. JVET Meeting, Jan. 6, 2021-Jan. 15, 2021, Teleconference, (The Joint Video Exploration Team of ISO… [cited by applicant]
Wang (Qualcomm) H., et al., “AHG11: Neural Network-Based In-Loop Filter Performance with No Deblocking Filtering Stage”, 21. JVET Meeting, Jan. 6, 2021-Jan. 15, 2021, Teleconference, (The Joint Video Exploration Team of… [cited by applicant]
Bross B., et al., “Versatile Video Coding (Draft 10)”, JVET-S2001-vH, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 19th Meeting: by teleconference, Jun. 22-Jul. 1, 2020, 551 Pages. [cited by applicant]
Bross B., et al., “Versatile Video Coding (Draft 9),” 130th MPEG Meeting, 18th JVET Meeting, Apr. 15, 2020-Apr. 24, 2020, Alpbach, (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11 and JVET of ITU-T SG 16 WP 3 and … [cited by applicant]
Chen J., et al., “AHG11: In-Loop Filtering with Convolutional Neural Network and Large Activation”, JVET-U0104-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21th Meeting: by teleconfer… [cited by applicant]
Hsiao Y-L., et al., “AHG9: Convolutional Neural Network Loop Filter”, JVET-M0159-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, pp. 1… [cited by applicant]
ITU-T H.265: “Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, High Efficiency Video Coding”, The International Telecommunication Union, Jun. 2019, 696 Pages. [cited by applicant]
Liu S., at al., “JVET Common Test Conditions and Evaluation Procedures for Neural Network-Based Video Coding Technology”, JVET-V2016-v3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISo/IEC JTC 1/SC 29, 22nd … [cited by applicant]
Liu S., et al., “JVET Common Test Conditions and Evaluation Procedures for Neural Network-Based Video Coding Technology”, JVET-T2006-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11,… [cited by applicant]
Liu S., et al., “JVET Common Test Conditions and Evaluation Procedures for Neural Network-Based Video Coding Technology”, JVET-U2016-r1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11,… [cited by applicant]
Timofte R., et al., “DIV2K Dataset: DIVerse 2K Resolution High Quality Images as Used for the Challenges @ NTIRE (CVPR 2017 (http://www.vision.ee.ethz.ch/ntire17) and CVPR 2018 (http://www.vision.ee.ethz.ch/ntire18)) an… [cited by applicant]
Wang H., et al., “AHG11: Neural Network-Based In-Loop Filter Performance with No Deblocking Filtering Stage”, JVET-U0115-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21st Meeting, by … [cited by applicant]
Wang H., et al., “EE: Tests on Neural Network-Based In-Loop Filter”, JVET-U0094-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 21st Meeting, by teleconference, Jan. 6-15, 2021, pp. 1-9. [cited by applicant]
Wang H., et al., “EE1-1.3: Test on Neural Network-based In-Loop Filter with No Deblocking Filtering Stage”, JVET-V0114-v3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, by te… [cited by applicant]
Wang H., et al., “EE1-1.4: Test on Neural Network-Based In-Loop Filter with Large Activation Layer”, JVET-V0115-v3, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, by teleconfe… [cited by applicant]
Wang H., et al., “EE1-1.4: Test on Neural Network-Based In-Loop Filter with Large Activation Layer”, JVET-W0130-v1, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 23rd Meeting, by teleconfe… [cited by applicant]
Wang H., et al., “EE1-Related: Neural Network-Based in-Loop Filter with Constrained Computational Complexity”, JVET-W0131-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 23rd Meeting, by… [cited by applicant]