IP Library Granted Patent US 11,636,343
Granted Patent B2
US 11,636,343 · App. 16/584,214 · Granted Apr 25, 2023

Systems and methods for neural network pruning with accuracy preservation

Inventor: Dan Alistarh (Meyrin, CH)
Assignee: Neuralmagic Inc.
G06N3/082G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,343
App. No.
16/584,214
Granted
Apr 25, 2023
Kind
B2
Abstract

Training a neural network (NN) may include training a NN N, and for S, a version of N to be sparsified (e.g. a copy of N), removing NN elements from S to create a sparsified version of S, and training S using outputs from N (e.g. “distillation”). A boosting or reintroduction phase may follow sparsification: training a NN may include for a trained NN N and S, a sparsified version of N, re-introducing NN elements previously removed from S, and training S using outputs from N. The boosting phase need not use a NN sparsified by “distillation.” Training and sparsification, or training and reintroduction, may be performed iteratively or over repetitions.

Claims (46)

1. A method for training a neural network (NN), the method comprising:

training a NN N; and

for S, a version of N to be sparsified:

removing NN elements from S to create a sparsified version of S; and

training S using outputs from N.

2. The method of claim 1 wherein the outputs from N comprise the outputs of a softmax layer.

3. The method of claim 1 wherein the NN elements comprise links.

4. The method of claim 1 comprising performing the removing and training operations over a series of iterations.

5. The method of claim 1 comprising training S using, for an input, outputs from N based on that input and the ground truth for that input.

6. The method of claim 1 comprising re-introducing NN elements previously removed from S, and training S using outputs from N.

7. The method of claim 1 , comprising training S using a loss generated from outputs from N.

8. A method for training a neural network (NN), the method comprising:

for a trained NN N and S, a sparsified version of N:

re-introducing NN elements previously removed from S; and

training S using outputs from N.

9. The method of claim 8 wherein the outputs from N comprise the outputs of a softmax layer.

10. The method of claim 8 wherein the NN elements comprise links.

11. The method of claim 8 comprising performing the re-introducing and training operations over a series of iterations.

12. The method of claim 8 comprising training S using, for an input, outputs from N based on that input and the ground truth for that input.

13. The method of claim 8 wherein S is sparsified by training S using outputs from N.

14. The method of claim 8 , comprising training S using a loss generated from outputs from N.

15. A system for training a neural network (NN), the system comprising:

a memory, and

a processor configured to:

train a NN N; and

for S, a version of N to be sparsified:

remove NN elements from S to create a sparsified version of S; and

train S using outputs from N.

16. The system of claim 15 wherein the outputs from N comprise the outputs of a softmax layer.

17. The system of claim 15 wherein the NN elements comprise links.

18. The system of claim 15 wherein the processor is configured to perform the removing and training operations over a series of iterations.

19. The system of claim 15 wherein the processor is configured to train S using, for an input, outputs from N based on that input and the ground truth for that input.

20. The system of claim 15 wherein the processor is configured to re-introduce NN elements previously removed from S, and training S using outputs from N.

21. The system of claim 15 , wherein the processor is configured to train S using a loss generated from outputs from N.

22. A system for training a neural network (NN), the system comprising:

a memory, and

a processor configured to:

for a trained NN N and S, a sparsified version of N:

re-introduce NN elements previously removed from S; and

train S using outputs from N.

23. The system of claim 22 wherein the outputs from N comprise the outputs of a softmax layer.

24. The system of claim 22 wherein the NN elements comprise links.

25. The system of claim 22 wherein the processor is configured to perform the re-introducing and training operations over a series of iterations.

26. The system of claim 22 wherein the processor is configured to train S using, for an input, outputs from N based on that input and the ground truth for that input.

27. The system of claim 22 wherein S is sparsified by training S using outputs from N.

28. The system of claim 22 , wherein the processor is configured to train S using a loss generated from outputs from N.

Assignments (3)
CHANGE OF NAME Recorded Mar 3, 2026
From: RED HAT, INC.
To: RED HAT, LLC
Reel/Frame 074913/0759 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2025
From: NEURALMAGIC, INC.
To: RED HAT, INC.
Reel/Frame 072278/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2019
From: ALISTARH, DAN
To: NEURALMAGIC INC.
Reel/Frame 051021/0405 →
Continuity (2)
Provisional Application 62739505 · Oct 1, 2018
Related Publication 20200104717A1 · Apr 2, 2020
Cited By (1)
US 12,675,265