IP Library Patent Application 18802235
Patent Application
App. No. 18/802,235

TRAINING A TARGET ACTIVATION SPARSITY IN A NEURAL NETWORK

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/802,235
Abstract

Techniques are described herein for a method of training a target activation sparsity in a neural network. The method includes obtaining a nonlinear portion of a plurality of neurons in a neural network. The neural network is trained to perform a target task. The method further includes substituting the nonlinear portion for a dynamic nonlinear portion in the plurality of neurons in the neural network. The dynamic nonlinear portion is trained to activate or deactivate one or more neurons of the plurality of neurons. The method further includes retraining the neural network using a first loss function that minimizes a loss of the target task and a second loss function that minimizes a number of active neurons.

Claims (46)

1 . A method comprising:

obtaining a nonlinear portion of a plurality of neurons in a neural network, wherein the neural network is trained to perform a target task;

substituting the nonlinear portion for a dynamic nonlinear portion in the plurality of neurons in the neural network, wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the plurality of neurons; and

retraining the neural network using a first loss function that minimizes a loss of the target task and a second loss function that minimizes a number of active neurons.

2 . The method of claim 1 , wherein the nonlinear portion of a neuron of the plurality of neurons is a preactivation distribution of the neuron.

3 . The method of claim 2 , wherein the preactivation distribution is based on a nonlinear activation function of the neuron of the plurality of neurons.

4 . The method of claim 2 , wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the neural network further comprises:

ordering samples of the preactivation distribution of the one or more neurons; and

selecting a number of neurons to activate responsive to a top number of the one or more neurons.

5 . The method of claim 1 , wherein obtaining the nonlinear portion of the plurality of neurons in the neural network further comprises:

computing a mean and a standard deviation of a preactivation distribution of a neuron of the plurality of neurons; and

determining a statistical model using the mean and the standard deviation.

6 . The method of claim 1 , wherein the retrained neural network is a sparse neural network having a target number of inactive neurons.

7 . The method of claim 1 , further comprising:

receiving a number of neurons in the neural network to be inactive, wherein the second loss function that minimizes the number of active neurons is based on the number of neurons in the neural network to be inactive.

8 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

obtaining a nonlinear portion of a plurality of neurons in a neural network, wherein the neural network is trained to perform a target task;

substituting the nonlinear portion for a dynamic nonlinear portion in the plurality of neurons in the neural network, wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the plurality of neurons; and

retraining the neural network using a first loss function that minimizes a loss of the target task and a second loss function that minimizes a number of active neurons.

9 . The non-transitory computer-readable medium of claim 8 , wherein the nonlinear portion of a neuron of the plurality of neurons is a preactivation distribution of the neuron.

10 . The non-transitory computer-readable medium of claim 9 , wherein the preactivation distribution is based on a nonlinear activation function of the neuron of the plurality of neurons.

11 . The non-transitory computer-readable medium of claim 9 , wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the neural network further comprises operations including:

ordering samples of the preactivation distribution of the one or more neurons; and

selecting a number of neurons to activate responsive to a top number of the one or more neurons.

12 . The non-transitory computer-readable medium of claim 8 , wherein obtaining the nonlinear portion of the plurality of neurons in the neural network further comprises operations including:

computing a mean and a standard deviation of a preactivation distribution of a neuron of the plurality of neurons; and

determining a statistical model using the mean and the standard deviation.

13 . The non-transitory computer-readable medium of claim 8 , wherein the retrained neural network is a sparse neural network having a target number of inactive neurons.

14 . The non-transitory computer-readable medium of claim 8 , wherein the operations further comprise:

receiving a number of neurons in the neural network to be inactive, wherein the second loss function that minimizes the number of active neurons is based on the number of neurons in the neural network to be inactive.

15 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

obtaining a nonlinear portion of a plurality of neurons in a neural network, wherein the neural network is trained to perform a target task;

substituting the nonlinear portion for a dynamic nonlinear portion in the plurality of neurons in the neural network, wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the plurality of neurons; and

retraining the neural network using a first loss function that minimizes a loss of the target task and a second loss function that minimizes a number of active neurons.

16 . The system of claim 15 , wherein the nonlinear portion of a neuron of the plurality of neurons is a preactivation distribution of the neuron.

17 . The system of claim 16 , wherein the dynamic nonlinear portion is trained to activate or deactivate one or more neurons in the neural network further comprises operations including:

ordering samples of the preactivation distribution of the one or more neurons; and

selecting a number of neurons to activate responsive to a top number of the one or more neurons.

18 . The system of claim 15 , wherein obtaining the nonlinear portion of the plurality of neurons in the neural network further comprises operations including:

computing a mean and a standard deviation of a preactivation distribution of a neuron of the plurality of neurons; and

determining a statistical model using the mean and the standard deviation.

19 . The system of claim 15 , wherein the retrained neural network is a sparse neural network having a target number of inactive neurons.

20 . The system of claim 15 , wherein the operations further comprise:

receiving a number of neurons in the neural network to be inactive, wherein the second loss function that minimizes the number of active neurons is based on the number of neurons in the neural network to be inactive.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2025
From: TENYX, INC.
To: SALESFORCE, INC.
Reel/Frame 070003/0174 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2024
From: KALAJDZIEVSKI, DAMJAN; COSENTINO, ROMAIN; SHEKKIZHAR, SARATH
To: TENYX, INC.
Reel/Frame 068266/0629 →