IP Library Granted Patent US 12,537,949
Granted Patent B2
US 12,537,949 · App. 17/621,476 · Granted Jan 27, 2026

Methods and apparatus for kernel tensor and tree partition based neural network compression framework

Inventors: Hua Yang (Plainsboro, NJ); Duanshun Li (Plainsboro, NJ); Dong Tian (Boxborough, MA); Yuwen He (San Diego, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/119H04N19/105H04N19/172H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,537,949
App. No.
17/621,476
Granted
Jan 27, 2026
Kind
B2
Abstract

A method of encoding or decoding a video comprising a current picture, a first reference picture, and a weight tensor associated with a trained neural network (NN) model are provided. The method includes generating any number of kernel tensors, input channels and output channels associated with the weight tensor, each kernel tensor being associated with any of: a layer type, an input signal type, and a tree partition type, and each kernel tensor including weight coefficients, generating, for each of the any number of kernel tensors, tree partitions for any of a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU) according to respective tree partition types associated with each of the any number of kernel tensors, and generating a compressed representation of the trained NN model by compressing and coding the any number of kernel tensors.

Claims (36)

1 . A method of decoding a trained neural network (NN) (TNN), the method comprising:

determining a first weight tensor (WT) associated with the TNN and a second WT associated with the TNN;

determining information associated with the TNN, wherein the information associated with the TNN indicates a layer type associated with a convolution NN (CNN);

generating an image based at least on an arrangement of the first WT and the second WT of the TNN;

generating a CNN layer based on the arrangement of the first WT and the second WT of the TNN, wherein the CNN layer includes a plurality of kernel tensors (KTs), wherein the plurality of the KTs includes a set of values of the first WT of the TNN and the second WT of the TNN, and wherein for each KT, an entirety of the each KT is used for generating tree partitions; and

outputting an NN model, wherein the NN model includes at least the generated CNN layer.

2 . The method of claim 1 , wherein, for the each KT, the entirety of the each KT is used for generating tree partitions for any of a prediction unit (PU) or a transform unit (TU).

3 . The method of claim 1 , further comprising using or selecting any of a coding syntax, a coding mode, or a coding method according to any of a dimension of a KT from the plurality of KTs and a size of the KT from the plurality of KTs.

4 . The method of claim 1 , wherein the layer type is any of: a convolutional layer type, a fully connected layer type, and a bias layer type.

5 . The method of claim 1 , wherein the method further comprises:

receiving a signal, wherein the signal is associated with one of a three-dimensional (3D) signal type associated with a video or a point cloud, a two-dimensional (2D) signal type associated with the image, or a one-dimensional (1D) signal type associated with audio, wherein the received signal is configured to be a picture and a number of reference pictures according to the arrangement of the received WTs of the TNN.

6 . The method of claim 1 , wherein the KTs are associated with a tree partition type (TPT), and wherein the TPT is one of a single tree type, a multi-tree type, or a mixed tree type.

7 . The method of claim 1 , wherein the plurality of KTs are divided into a plurality of sub-tenors, wherein each sub-tensor in the plurality of sub-tensors includes a respective number of weight coefficients.

8 . The method of claim 7 , further comprising compressing coding of the plurality of KTs by compressing and coding the plurality of sub-tensors associated with the KTs.

9 . The method of claim 8 , further comprising compressing and coding at least one of the KTs or sub-tensors by performing tree partitions on the at least one of the KTs or the sub-tensors and a quantization unit (QU).

10 . The method of claim 1 , wherein generating the CNN layer further comprises,

determining, for the plurality of KTs, at least one of: an input signal type of the each KT in the plurality of KTs, or a plurality of tree partition types (TPTs) of the each KT; and

decoding the plurality of KTs according to one of the input signal type of the each KT or the determined TPTs for each KT in the plurality of KTs, wherein the TPTs are one of a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

11 . A wireless transmit receive unit (WTRU) comprising circuitry including a transmitter, a receiver, a processor and memory, configured to:

determining a first weight tensor (WT) associated with the TNN and a second WT associated with the TNN;

determine information associated with the TNN, wherein the information associated with the TNN indicates a layer type associated with a convolution NN (CNN);

generate an image based at least on an arrangement of the first WT and the second WT of the TNN;

generate a CNN layer based on the arrangement of the first WT and the second WT of the TNN, wherein the CNN layer includes a plurality of kernel tensors (KTs), wherein the plurality of the KTs includes a set of values of the first WT of the TNN and the second WT of the TNN, and wherein for each KT, an entirety of the each KT is used for generating tree partitions; and

output an NN model, wherein the NN model includes at least the generated CNN layer.

12 . The WTRU of claim 11 , wherein, for the each KT, the entirety of the each KT is used for generating tree partitions for any of a prediction unit (PU) or a transform unit (TU).

13 . The WTRU of claim 11 , configured to use or select any of a coding syntax, a coding mode, or a coding method according to any of a dimension of a KT from the plurality of KTs and a size of the KT from the plurality of KTs.

14 . The WTRU of claim 11 , wherein the layer type is any of: a convolutional layer type, a fully connected layer type, and a bias layer type.

15 . The WTRU of claim 11 , wherein the processor is further configured to:

receive a signal, wherein the signal is associated with one of a three-dimensional (3D) signal type associated with a video or a point cloud, a two-dimensional (2D) signal type associated with the image, or a one-dimensional (1D) signal type associated with audio, wherein the received signal is configured to be a picture and a number of reference pictures according to the arrangement of the received WTs of the TNN.

16 . The WTRU of claim 11 , wherein the KTs are associated with a tree partition type (TPT), and wherein the TPT is one of a single tree type, a multi-tree type, or a mixed tree type.

17 . The WTRU of claim 11 , wherein the plurality of KTs are divided into a plurality of sub-tenors, wherein each sub-tensor in the plurality of sub-tensors includes a respective number of weight coefficients.

18 . The WTRU of claim 17 , wherein the processor is further configured to compress coding of the plurality of KTs by compressing and coding the plurality of sub-tensors associated with the KTs.

19 . The WTRU of claim 18 , wherein the processor is further configured to compress and code at least one of the KTs or sub-tensors by performing tree partitions on the at least one of the KTs or the sub-tensors and a quantization unit (QU).

20 . The WTRU of claim 1 , wherein for the generation of the CNN layer, the processor is further configured to:

determine, for the plurality of KTs each KT, at least one of: an input signal type of the each KT in the plurality of KTs, or a plurality of tree partition types (TPTs) of the each KT; and

decode the plurality of KTs according to one of the input signal type of the each KT or the determined TPTs for each KT in the plurality of KTs, wherein the TPTs are one of a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2022
From: YANG, HUA; LI, DUANSHUN; TIAN, DONG; HE, YUWEN
To: VID SCALE, INC.
Reel/Frame 058649/0696 →
Continuity (2)
Provisional Application 62869679 · Jul 2, 2019
Related Publication 20220360778A1 · Nov 10, 2022
References Cited (34)
US 20130107950A1 · Guo et al. · 2013 [cited by applicant]
US 20160044314A1 · Rinaldi · 2016 [cited by applicant]
US 20190035113A1 · Salvi et al. · 2019 [cited by applicant]
US 20190075301A1 · Chou et al. · 2019 [cited by applicant]
US 20210125070A1 · Wang · 2021 [cited by examiner]
CN 106663085A · 2017 [cited by applicant]
WO 2019086104A1 · 2019 [cited by applicant]
Wang et al., “Huawei's response to the Call for Proposal on Neural Network Compression”, 126. MPEG Meeting; Mar. 25, 2019-Mar. 29, 2019; Geneva, No. m47491, Mar. 29, 2019. (Year: 2019). [cited by examiner]
“Description of Core Experiments on Compression of Neural Networks for Multimedia Content Description and Analysis”, ISO/IEC JTC1/SC29/WG11/N18461, MPEG Meeting 126, Geneva, XP0302308730, Mar. 25-29, 2019, 17 pages (Yea… [cited by examiner]
Agustsson et al., “Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations”, In Advances in Neural Information Processing Systems, arXiv:1704.00648v2, Jun. 8, 2017, pp. 1-16. [cited by applicant]
Aytekin et al., “Compressibility Loss for Neural Network Weights”, arXiv:1905.01044v1, May 3, 2019, 7 pages. [cited by applicant]
Bross et al., “High Efficiency Video Coding (HEVC) Text Specification Draft 10 (for FDIS & Consent)”, JCTVC-L1003_V1, Editor, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29… [cited by applicant]
Chen et al., “Algorithm Description for Versatile Video Coding and Test Model 5 (VTM 5)”, JVET-N1002-v1, Editors, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting: Geneva, … [cited by applicant]
Choi et al., “Towards the Limit of Network Quantization”, ICLR 2017, arXiv:1612.01543v2, Nov. 13, 2017, pp. 1-14. [cited by applicant]
Choi et al., “Universal Deep Neural Network Compression”, arXiv:1802.02271v2, Feb. 21, 2019, 5 pages. [cited by applicant]
Denil et al., “Predicting Parameters in Deep Learning”, In Advances in Neural Information Processing Systems, Oct. 27, 2013, pp. 1-9. [cited by applicant]
Denton et al., “Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation”, In Advances in Neural Information Processing Systems, 2014, pp. 1-9. [cited by applicant]
Duanshun et al., “NNR CE2-related: On Quantization for Neural Network Compression”, ISO/IEC JTC1/SC29/WG11, No. m49408, 127 MPEG Meeting, Jul. 8-12, 2019. [cited by applicant]
Gong et al., “Compressing Deep Convolutional Networks Using Vector Quantization”, ICLR 2015, arXiv:1412.6115v1, Dec. 18, 2014, pp. 1-10. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, ICLR 2016, arXiv:1510.00149v5, Feb. 15, 2016, pp. 1-14. [cited by applicant]
ISO/IEC, “Updated Call for Proposals on Neural Network Compression”, MPEG Requirements, ISO/IEC JTC1/SC29/WG11/N18162, Marrakech, MA, Jan. 2019, 8 pages. [cited by applicant]
ISO/IEC, “Updated Evaluation Framework for Compressed Representation of Neural Networks”, Requirements Subgroup, ISO/IEC JTC1/SC29/WG11/N18129, Marrakech, MA, Jan. 2019, 12 pages. [cited by applicant]
Jain et al., “Response to the Call for Proposals on Neural Network Compression, Low Displacement Rank based compression of Deep Neural Networks”, Technicolor, ISO/IEC JTC1/SC29/WG11 MPEG2019/M47493, Geneva, Switzerland,… [cited by applicant]
Jin et al., “Flattened Convolutional Neural Networks for Feedforward Acceleration”, ICLR 2015, arXiv:1412.5474v4, Nov. 20, 2015, pp. 1-11. [cited by applicant]
Kim et al., “Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications”, arXiv:1511.06530v2, Feb. 24, 2016, pp. 1-16. [cited by applicant]
Laude et al., “Neural Network Compression Using Transform Coding and Clustering”, arXiv:1805.07258v1, May 18, 2018, 4 pages. [cited by applicant]
Li et al., “A Deep Convolutional Neural Network Approach for Complexity Reduction on Intra-Mode HEVC”, Proceedings of the IEEE International Conference on Multimedia and Expo (ICME) 2017, Jul. 10-14, 2017, pp. 1255-1260. [cited by applicant]
Polino et al., “Model Compression via Distillation and Quantization”, ICLR 2018, arXiv:1802.05668v1, Feb. 15, 2018, pp. 1-21. [cited by applicant]
Sullivan et al., “Overview of the High Efficiency Video Coding (HEVC) Standard”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, pp. 1649-1668. [cited by applicant]
Wiedemann et al., “Entropy-Constrained Training of Deep Neural Networks”, arXiv:1812.07520v2, Dec. 19, 2018, 8 pages. [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 5)”, JVET-N1001-V10, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 14th Meeting, Geneva, Switzerland, Mar. 19, 2019, 406 pages. [cited by applicant]
Bossen, Frank, “VTM Software Manual”, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 34 pages. [cited by applicant]
ISO/IEC, , “Description of Core Experiments on Compression of Neural Networks for Multimedia Content Description and Analysis”, ISO/IEC JTC1/SC29/WG11/N18461, MPEG Meeting 126, Geneva, XP0302308730, Mar. 25-29, 2019, 17… [cited by applicant]
ISO/IEC, “Evaluation Results of the Call for Proposals on Neural Network Compression”, ISO/IEC JTC1/SC29/WG11, No. n18352, 126 MPEG Meeting, XP030208622, Mar. 25-Mar. 29, 2019, 15 pages. [cited by applicant]