IP Library Granted Patent US 12,456,076
Granted Patent B2
US 12,456,076 · App. 18/537,689 · Granted Oct 28, 2025

Activation based dynamic network pruning

Inventor: Minhoo Kang (Seongnam-si, KR)
Assignee: REBELLIONS INC.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,076
App. No.
18/537,689
Granted
Oct 28, 2025
Kind
B2
Abstract

A method for controlling operations of a machine learning model is performed by one or more processors and includes determining a threshold for skipping an operation, acquiring an activation value associated with at least one layer included in the machine learning model, determining whether the activation value is less than the threshold, and if the activation value is less than the threshold, controlling the operations of the machine learning model such that an operation associated with the activation value is skipped in the machine learning model.

Claims (45)

1. A method for controlling operations of a machine learning model, the method being performed by one or more processors and comprising:

receive a first command to statistically determine a threshold for skipping an operation associated with a node and continuously maintain the threshold;

receive a second command to dynamically determine the threshold for skipping the operation associated with the node;

dynamically determining the threshold for skipping the operation associated with the node based on a distribution of a plurality of activation values associated with at least one layer included in the machine learning model, the machine learning model comprising a plurality of layers including an input layer that receives an input signal and an output layer that outputs an output signal and a plurality of hidden layers positioned between the input layer and the output layer to receive a signal from the input layer, extract features, and transmit the features to the output layer, wherein the plurality of layers includes a first layer including a first set of nodes and a second layer including a second set of nodes;

acquiring an activation value associated with at least one layer included in the machine learning model;

determining whether the activation value is less than the threshold; and

if the activation value is less than the threshold, controlling the operations of the machine learning model such that an operation associated with the activation value is skipped in the machine learning model,

wherein the activation value includes an output value from the first layer included in the machine learning model, and

the controlling the operations of the machine learning model includes, if operations associated with the second layer into which an output value from the first layer is input are performed and the output value from the first layer is less than the threshold, controlling the operations of the machine learning model by transmitting a skip command associated with the output value and storing the output value in a memory without changing the output value to zero such that an operation associated with the output value from the first layer is skipped in the second layer, performing the machine learning model by performing operations associated with the second set of nodes of the second layer and associated with values unassociated with the skip command and skipping the operations associated with the second set of nodes of the second layer and associated with the output value associated with the skip command regardless of whether the output value stored in the memory is zero;

wherein the determining the threshold includes:

when the activation value is less than the threshold, decreasing the threshold by a predetermined amount;

increasing a counter value indicating a time during which the threshold does not change when it is determined that the activation value is equal to or greater than the threshold; and

if the counter value reaches a predetermined value, resetting the counter value and updating the threshold such that the threshold is increased by the predetermined amount.

2. The method according to claim 1 , wherein the activation value is expressed as a floating-point number in a predetermined format,

wherein the method further comprising:

acquiring an exponent from the activation value expressed as the floating-point number;

determining whether the acquired exponent is less than a threshold; and

if the acquired exponent is less than the threshold, controlling the operations of the machine learning model such that an operation associated with the activation value is skipped in the machine learning model.

3. The method according to claim 2 , wherein the threshold is a value obtained by multiplying 2 by a predetermined number of times.

4. The method according to claim 2 , wherein each of the activation value and the threshold is a binary number.

5. The method according to claim 2 , further comprising, before the acquiring the activation value, determining the threshold based on a distribution of a plurality of activation values associated with the at least one layer.

6. The method according to claim 5 , wherein the determining the threshold includes:

determining a target activation value for comparison from a plurality of activation values associated with the at least one layer;

acquiring a target exponent from the determined target activation value; and

if the acquired target exponent is less than the threshold, updating the threshold such that the threshold is decreased by a predetermined amount.

7. The method according to claim 6 , wherein the updating the threshold includes decreasing the threshold by dividing the threshold by 2.

8. The method according to claim 1 , wherein the updating the threshold includes updating the threshold by multiplying the threshold by 2.

9. The method according to claim 2 , wherein the acquiring the activation value includes acquiring the activation value from a memory, wherein the activation value includes an output value from a first layer included in the machine learning model, and

the controlling the operations of the machine learning model includes,

if operations associated with a second layer into which the output value from the first layer is input are performed, and an exponent of the output value from the first layer is less than the threshold, controlling the operations of the machine learning model such that an operation associated with the activation value is skipped.

10. A non-transitory computer-readable recording medium storing instructions for execution by the one or more processors that, when executed by the one or more processors, cause the one or more processors to perform the method according to claim 1 .

11. A processing system comprising: a memory storing one or more instructions; and a processor configured to execute one or more instructions in the memory to:

receive a first command to statistically determine a threshold for skipping an operation associated with a node and continuously maintain the threshold;

receive a second command to dynamically determine the threshold for skipping the operation associated with the node;

dynamically determine the threshold for skipping the operation associated with the node based on a distribution of a plurality of activation values associated with at least one layer included in a machine learning model, the machine learning model comprising a plurality of layers including an input layer that receives an input signal and an output layer that outputs an output signal and a plurality of hidden layers positioned between the input layer and the output layer to receive a signal from the input layer, extract features, and transmit the features to the output layer, wherein the plurality of layers includes a first layer including a first set of nodes and a second layer including a second set of nodes;

acquire an activation value associated with at least one layer included in a machine learning model;

determine whether the activation value is less than the threshold; and

if the activation value is less than the threshold, control operations of the machine learning model such that an operation associated with the activation value is skipped in the machine learning model,

wherein the activation value includes an output value from the first layer included in the machine learning model, and

the processor is further configured to, if operations associated with the second layer into which an output value from the first layer is input are performed, and the output value from the first layer is less than the threshold, control the operations of the machine learning model by transmitting a skip command associated with the output value and storing the output value in a memory without changing the output value to zero such that an operation associated with the output value from the first layer is skipped in the second layer,

performing the machine learning model by performing operations associated with the second set of nodes of the second layer and associated with values unassociated with the skip command and skipping the operations associated with the second set of nodes of the second layer and associated with the output value associated with the skip command regardless of whether the output value stored in the memory is zero;

wherein the determining the threshold includes:

when the activation value is less than the threshold, decreasing the threshold by a predetermined amount;

increasing a counter value indicating a time during which the threshold does not change when it is determined that the activation value is equal to or greater than the threshold; and

if the counter value reaches a predetermined value, resetting the counter value and updating the threshold such that the threshold is increased by the predetermined amount.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded May 22, 2025
From: REBELLIONS INC.; SAPEON KOREA INC.
To: REBELLIONS INC.
Reel/Frame 071349/0150 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2023
From: KANG, MINHOO
To: REBELLIONS INC.
Reel/Frame 065850/0537 →
Priority Claims (1)
KR 10-2023-0068299 · May 26, 2023 · national
Continuity (1)
Related Publication 20240394594A1 · Nov 28, 2024
References Cited (32)
US 20180232640A1 · Ji · 2018 [cited by examiner]
US 20190057308A1 · Cho · 2019 [cited by examiner]
US 20190065195A1 · Pool · 2019 [cited by examiner]
US 20190228344A1 · Hong et al. · 2019 [cited by applicant]
US 20190362235A1 · Xu · 2019 [cited by examiner]
US 20200134459A1 · Zeng · 2020 [cited by examiner]
US 20200242454A1 · Kim et al. · 2020 [cited by applicant]
US 20210004442A1 · Sapugay · 2021 [cited by examiner]
US 20220019897A1 · Park et al. · 2022 [cited by applicant]
US 20220245457A1 · Srinivas · 2022 [cited by examiner]
US 20220327367A1 · Judd · 2022 [cited by examiner]
US 20230153623A1 · Bond · 2023 [cited by examiner]
US 20240176981A1 · Khodamoradi · 2024 [cited by examiner]
KR 1020190088643A · 2019 [cited by applicant]
KR 1020200094534A · 2020 [cited by applicant]
KR 1020210125911A · 2021 [cited by applicant]
KR 1020220010383A · 2022 [cited by applicant]
KR 1020220082220A · 2022 [cited by applicant]
KR 1020220133566A · 2022 [cited by applicant]
Sim et al., “SparTANN: Sparse training accelerator for neural networks with threshold-based sparsification”, 2020, Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design, vol. 2020, pp. … [cited by examiner]
Vestias et al., “Fast Convolutional Neural Networks in Low Density FPGAs Using Zero-Skipping and Weight Pruning”, 2019, Electronics, vol. 8, pp. 1-24 (Year: 2019). [cited by examiner]
Whatmough et al., “DNN Engine: A 28-nm Timing-Error Tolerant Sparse Deep Neural Network Processor for IoT Applications”, 2018, IEEE Journal of Solid-State Circuits, vol. 53 No. 9, pp. 2722-2731 (Year: 2018). [cited by examiner]
Bakhtiary et al., “Winner takes all hashing for speeding up the training of neural networks in large class problems”, 2017, Pattern Recognition Letters, vol. 93, pp. 38-47 (Year: 2017). [cited by examiner]
Wu et al., “Guessing Outputs of Dynamically Pruned CNNs Using Memory Access Patterns”, 2021, IEEE Computer Architecture Letters, vol. 20 No. 2, pp. 98-101 (Year: 2021). [cited by examiner]
Kurtz et al., “Inducing and exploiting activation sparsity for fast inference on deep neural networks”, 2020, International Conference on Machine Learning, vol. 2020, pp. 5533-5543 (Year: 2020). [cited by examiner]
Raihan et al., “Sparse weight activation training”, 2020, Advances in Neural Information Processing Systems, vol. 33 (2020), pp. 15625-15638 (Year: 2020). [cited by examiner]
Machupalli et al., “Review of ASIC accelerators for deep neural network”, 2022, Microprocessors and Microsystems, vol. 89, pp. 1-19 (Year: 2022). [cited by examiner]
Im et al., “DT-CNN: An Energy-Efficient Dilated and Transposed Convolutional Neural Network Processor for Region of Interest Based Image Segmentation”, 2020, IEEE Transactions on Circuits and Systems I: Regular Papers, … [cited by examiner]
Huan et al., “A Low-Power Accelerator for Deep Neural Networks with Enlarged Near-Zero Sparsity”, 2017, arXiv, v1, pp. 1-5 (Year: 2017). [cited by examiner]
Lee et al., “An accurate and fair evaluation methodology for SNN-based inferencing with full-stack hardware design space explorations”, 2021, Neurocomputing, vol. 455, pp. 125-138 (Year: 2021). [cited by examiner]
Zhang et al., “Adaptive Filter Pruning via Sensitivity Feedback”, Mar. 8, 2023, IEEE Transactions on Neural Networks and Learning Systems, vol. 35 No. 8, pp. 10996-11008 (Year: 2023). [cited by examiner]
Haberer et al., “Activation Sparsity and Dynamic Pruning for Split Computing in Edge AI”, 2022, DistributedML '22: Proceedings of the 3rd International Workshop on Distributed Machine Learning, vol. 3 (2022), pp. 30-36 … [cited by examiner]