IP Library Granted Patent US 11,537,892
Granted Patent B2
US 11,537,892 · App. 16/632,215 · Granted Dec 27, 2022

Slimming of neural networks in machine learning environments

Inventors: Shoumeng Yan (Beijing, CN); Jianguo Li (Beijing, CN); Zhuang Liu (Beijing, CN)
Assignee: INTEL CORPORATION
G06N3/082G06N3/0454G06N3/063G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,537,892
App. No.
16/632,215
Granted
Dec 27, 2022
Kind
B2
Abstract

A mechanism is described for facilitating slimming of neural networks in machine learning environments. A method of embodiments, as described herein, includes learning a first neural network associated with machine learning processes to be performed by a processor of a computing device, where learning includes analyzing a plurality of channels associated with one or more layers of the first neural network. The method may further include computing a plurality of scaling factors to be associated with the plurality of channels such that each channel is assigned a scaling factor, wherein each scaling factor to indicate relevance of a corresponding channel within the first neural network. The method may further include pruning the first neural network into a second neural network by removing one or more channels of the plurality of channels having low relevance as indicated by one or more scaling factors of the plurality of scaling factors assigned to the one or more channels.

Claims (27)

1. An apparatus comprising:

one or more processors to:

learn a first neural network associated with machine learning processes, wherein learning includes analyzing a plurality of channels associated with one or more layers of the first neural network;

compute a plurality of scaling factors to be associated with the plurality of channels such that each channel is assigned a scaling factor, wherein each scaling factor to indicate relevance of a corresponding channel within the first neural network; and

prune the first neural network into a second neural network by removing one or more channels of the plurality of channels having low relevance as indicated by one or more scaling factors of the plurality of scaling factors assigned to the one or more channels, wherein the one or more scaling factors comprise one or more numbers to indicate one or more relevance levels of the plurality of channels within the first neural network such that a minimum relevance level is indicated by a minimum threshold number.

2. The apparatus of claim 1 , wherein the one or more processors are further to train or fine tune the second neural network into a third neural network, wherein the third neural network is without the removed one or more channels such that the third neural network having few channels than the plurality of channels demands fewer processing resources than the first neural network.

3. The apparatus of claim 2 , wherein the first neural network comprises an initial wide neural network, wherein the second neural network comprises a pruned neural network, and wherein the third neural network comprises a final slim neural network.

4. The apparatus of claim 1 , wherein the one or more processors are further to detect the first neural network and observe importance of each of the plurality of channels within the first neural network, wherein the importance of a channel indicates whether the first neural network is sustainable without the channel.

5. The apparatus of claim 1 , wherein the removed one or more channels are assigned one or more scaling factors having one or more numbers equal to or lower than the minimum threshold number.

6. The apparatus of claim 1 , wherein the one or more processors comprises a graphics processor co-located with an application processor on a common semiconductor package.

7. A method comprising:

learning a first neural network associated with machine learning processes to be performed by a processor of a computing device, wherein learning includes analyzing a plurality of channels associated with one or more layers of the first neural network;

computing a plurality of scaling factors to be associated with the plurality of channels such that each channel is assigned a scaling factor, wherein each scaling factor to indicate relevance of a corresponding channel within the first neural network; and

pruning the first neural network into a second neural network by removing one or more channels of the plurality of channels having low relevance as indicated by one or more scaling factors of the plurality of scaling factors assigned to the one or more channels, wherein the one or more scaling factors comprise one or more numbers to indicate one or more relevance levels of the plurality of channels within the first neural network such that a minimum relevance level is indicated by a minimum threshold number.

8. The method of claim 7 , further comprising training or fine-tuning the second neural network into a third neural network, wherein the third neural network is without the removed one or more channels such that the third neural network having few channels than the plurality of channels demands fewer processing resources than the first neural network.

9. The method of claim 8 , wherein the first neural network comprises an initial wide neural network, wherein the second neural network comprises a pruned neural network, and wherein the third neural network comprises a final slim neural network.

10. The method of claim 7 , further comprising detecting the first neural network and observe importance of each of the plurality of channels within the first neural network, wherein the importance of a channel indicates whether the first neural network is sustainable without the channel.

11. The method of claim 7 , wherein the removed one or more channels are assigned one or more scaling factors having one or more numbers equal to or lower than the minimum threshold number.

12. The method of claim 7 , wherein the processor comprises a graphics processor co-located with an application processor on a common semiconductor package.

13. At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:

learning a first neural network associated with machine learning processes to be performed by a processor of the computing device, wherein learning includes analyzing a plurality of channels associated with one or more layers of the first neural network;

computing a plurality of scaling factors to be associated with the plurality of channels such that each channel is assigned a scaling factor, wherein each scaling factor to indicate relevance of a corresponding channel within the first neural network; and

pruning the first neural network into a second neural network by removing one or more channels of the plurality of channels having low relevance as indicated by one or more scaling factors of the plurality of scaling factors assigned to the one or more channels, wherein the one or more scaling factors comprise one or more numbers to indicate one or more relevance levels of the plurality of channels within the first neural network such that a minimum relevance level is indicated by a minimum threshold number.

14. The non-transitory machine-readable medium of claim 13 , wherein the operations further comprise training or fine-tuning the second neural network into a third neural network, wherein the third neural network is without the removed one or more channels such that the third neural network having few channels than the plurality of channels demands fewer processing resources than the first neural network.

15. The non-transitory machine-readable medium of claim 13 , wherein the first neural network comprises an initial wide neural network, wherein the second neural network comprises a pruned neural network, and wherein the third neural network comprises a final slim neural network.

16. The non-transitory machine-readable medium of claim 13 , wherein the operations further comprise detecting the first neural network and observe importance of each of the plurality of channels within the first neural network, wherein the importance of a channel indicates whether the first neural network is sustainable without the channel.

17. The non-transitory machine-readable medium of claim 13 , wherein the removed one or more channels are assigned one or more scaling factors having one or more numbers equal to or lower than the minimum threshold number, wherein the processor comprises a graphics processor co-located with an application processor on a common semiconductor package.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2020
From: LIU, ZHUANG; YAN, SHOUMENG; LI, JIANGUO
To: INTEL CORPORATION
Reel/Frame 051822/0488 →
Continuity (1)
Related Publication 20200234130A1 · Jul 23, 2020
Cited By (1)
US 12,374,031