IP Library Granted Patent US 12,340,313
Granted Patent B2
US 12,340,313 · App. 17/107,046 · Granted Jun 24, 2025

Neural network pruning method and system via layerwise analysis

Inventors: Enxu Yan (Los Altos, CA); Dongkuan Xu (Los Altos, CA); Jiachao Liu (Los Altos, CA)
Assignee: MOFFETT INTERNATIONAL CO., LIMITED
G06N3/082G06N3/045G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,313
App. No.
17/107,046
Granted
Jun 24, 2025
Kind
B2
Abstract

Embodiments disclosed herein allowed neural networks to be pruned. The inputs and outputs generated by a reference neural network are used to prune the reference neural network. The pruned neural network may have a subset of the weights that are in the reference neural network.

Claims (55)

1. A method, comprising:

obtaining a first neural network, wherein:

the first neural network is trained using a first set of training data;

the first neural network comprises a first set of nodes; and

the first neural network comprises a first set of connections that interconnect the first set of nodes;

generating a second neural network based on the first neural network, wherein generating the second neural network comprises

analyzing a set of intermediate layers of the first neural network,

determining subsets of weights for each intermediate layer of the intermediate layers, including generating, for each intermediate layer, a respective filter based on an input provided to the intermediate layer and a reference output, wherein the respective filter comprises a respective one of the subsets of weights, and

generating the second neural network based on the subsets of weights for each intermediate layer, wherein:

the second neural network comprises a second set of connections interconnecting a second set of nodes and the second set of connections comprises a subset of the first set of connections; and

the second neural network is generated without using the first set of training data.

2. The method of claim 1 , wherein the input is generated by:

stacking a first set of input feature maps to generate a first combined feature map, wherein the input comprises the first combined feature map.

3. The method of claim 2 , wherein the first set of input feature maps are generated by a first filter of the first neural network.

4. The method of claim 1 , wherein the reference output is generated by:

flattening an output feature map to generate a vector, wherein the reference output comprises the vector.

5. The method of claim 4 , wherein the output feature map is generated based on a second filter of the first neural network.

6. The method of claim 1 , wherein the subsets of weights for each intermediate layer are determined simultaneously for each intermediate layer.

7. The method of claim 1 , wherein generating the second neural network comprises:

generating the second neural network without training the second neural network.

8. The method of claim 1 , wherein the second set of nodes comprises a subset of the first set of nodes.

9. A method, comprising:

obtaining a first neural network, wherein:

the first neural network is trained by passing a set of training data through the first neural network a first number of times;

the first neural network comprises a first set of nodes; and

the first neural network comprises a first set of connections that interconnect the first set of nodes;

generating a second neural network based on the first neural network, wherein generating the second neural network comprises

analyzing a set of intermediate layers of the first neural network,

determining subsets of weights for each intermediate layer of the intermediate layers, including generating, for each intermediate layer, a respective filter based on an input provided to the intermediate layer and a reference output, wherein the respective filter comprises a respective one of the subsets of weights, and

generating the second neural network based on the subsets of weights for each intermediate layer, wherein:

the second neural network comprises a second set of connections interconnecting a second set of nodes and the second set of connections comprises a subset of the first set of connections; and

the second neural network is generated by passing the set of training data through the second neural network a second number of times, wherein the second number of times is smaller than the first number of times.

10. The method of claim 9 , wherein the input is generated by:

stacking a first set of input feature maps to generate a first combined feature map, wherein the input comprises the first combined feature map.

11. The method of claim 10 , wherein the first set of input feature maps are generated by a previous layer of the second neural network.

12. The method of claim 9 , wherein the reference output is generated by:

flattening an output feature map to generate a vector, wherein the reference output comprises the vector.

13. The method of claim 12 , wherein the output feature map is generated based on a second filter of the first neural network.

14. The method of claim 9 , wherein the subsets of weights for each intermediate layer are determine sequentially, layer by layer.

15. The method of claim 9 , wherein the second set of nodes comprises a subset of the first set of nodes.

16. An apparatus, comprising:

a memory configured to store data;

a processor coupled to the memory, the processor configured to:

obtain a first neural network, wherein:

the first neural network is trained using a first set of training data;

the first neural network comprises a first set of nodes; and

the first neural network comprises a first set of connections that interconnect the first set of nodes;

generate a second neural network based on the first neural network, wherein generating the second neural network comprises

analyzing a set of intermediate layers of the first neural network,

determining subsets of weights for each intermediate layer of the intermediate layers, including generating, for each intermediate layer, a respective filter based on an input provided to the intermediate layer and a reference output, wherein the respective filter comprises a respective one of the subsets of weights, and

generating the second neural network based on the subsets of weights for each intermediate layer, wherein:

the second neural network comprises a second set of connections interconnecting a second set of nodes and the second set of connections comprises a subset of the first set of connections; and

the second neural network is generated without using the first set of training data.

17. The apparatus of claim 16 , wherein the input is generated by:

stacking a first set of input feature maps to generate a first combined feature map, wherein the input comprises the first combined feature map.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: MOFFETT TECHNOLOGIES CO., LIMITED
To: MOFFETT INTERNATIONAL CO., LIMITED
Reel/Frame 061241/0626 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2020
From: YAN, ENXU; XU, DONGKUAN; LIU, JIACHAO
To: MOFFETT TECHNOLOGIES CO., LIMITED
Reel/Frame 054673/0565 →
Continuity (1)
Related Publication 20220172059A1 · Jun 2, 2022
References Cited (20)
US 11995551B2 · Koivisto · 2024 [cited by examiner]
US 20160293167A1 · Chen et al. · 2016 [cited by applicant]
US 20180336468A1 · Kadav · 2018 [cited by examiner]
US 20190108436A1 · David · 2019 [cited by examiner]
US 20190122394A1 · Shen · 2019 [cited by examiner]
US 20190362235A1 · Xu et al. · 2019 [cited by applicant]
US 20200160185A1 · Praveen · 2020 [cited by examiner]
US 20200311552A1 · A · 2020 [cited by examiner]
US 20200356860A1 · Kim · 2020 [cited by examiner]
US 20200364572A1 · Senn · 2020 [cited by examiner]
US 20210174175A1 · Choudhury · 2021 [cited by examiner]
CN 111210016A · 2020 [cited by applicant]
CN 111461324A · 2020 [cited by applicant]
CN 111738401A · 2020 [cited by applicant]
CN 111931901A · 2020 [cited by applicant]
JP 2018129033A · 2018 [cited by applicant]
TW 202042559A · 2020 [cited by applicant]
Lin et al. (NPL “Toward Compact ConvNets via Structure-Sparsity Regularized Filter Pruning”, vol. 31, No. 2, Feb. 2020) (Year: 2020). [cited by examiner]
Jian-Hao Luo et al. (NPL, “ThiNet: Pruning CNN Filters for a Thinner Net”, vol. 41, No. 10, Oct. 2019) (Year: 2019). [cited by examiner]
He et al. (NPL “Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks”, arXiv Aug. 21, 2018) (Year: 2018). [cited by examiner]
Cited By (1)
US 12,614,255