IP Library Granted Patent US 12670383
Granted Patent B2
US 12670383 · App. 17/548,692 · Granted Jun 30, 2026

System and method of using fractional adaptive linear unit as activation in artificial neural networks

Inventors: Anthony Daniel Rhodes (Portland, OR); Julio Cesar Zamora Esquivel (West Sacramento, CA); Lama Nachman (Santa Clara, CA)
Assignee: Intel Corporation
G06N3/08G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670383
App. No.
17/548,692
Granted
Jun 30, 2026
Kind
B2
Abstract

An apparatus is provided for deep learning. The apparatus accesses a neural network including an input layer, hidden layers, and an output layer. The apparatus adds an activation function to one or more of the hidden layers of the hidden layers and output layer. The activation function includes a tunable parameter, the value of which can be adjusted during the training of the neural network. The apparatus trains the neural network by inputting training samples into the neural network and determining internal parameters of the neural network based on the training samples. Determining the internal parameters includes determining a value of the tunable parameter based on the training samples. The apparatus may determine two different values of the tunable parameter for two different layers. The activation function may include another tunable parameter. The apparatus can determine a value for the other tunable parameter during the training of the neural network.

Claims (69)

1 . A method for deep learning based on a trainable activation function, the method comprising:

accessing a neural network comprising an input layer, a plurality of hidden layers, and an output layer;

adding the trainable activation function with a differentiation operator to one or more hidden layers of the plurality of hidden layers, the trainable activation function comprising a tunable parameter, the tunable parameter representing a fractional derivative order of the differentiation operator in the trainable activation function; and

training the neural network by inputting a plurality of training samples into the neural network, wherein training the neural network comprises:

determining a value of the tunable parameter based on the plurality of training samples.

2 . The method of claim 1 , wherein determining the value of the tunable parameter based on the plurality of training samples comprises:

confining a domain of the tunable parameter to a range from 0 to 2; and

determining the value of the tunable parameter within the domain based on the plurality of training samples.

3 . The method of claim 1 , wherein adding the trainable activation function with the differentiation operator to the one or more hidden layers of the plurality of hidden layers comprises:

adding the trainable activation function with the differentiation operator to a first hidden layer and a second hidden layer,

wherein training the neural network comprises:

determining a first value of the tunable parameter for the first hidden layer based on the plurality of training samples, and

determining a second value of the tunable parameter for the second hidden layer based on the plurality of training samples.

4 . The method of claim 3 , wherein the first value of the tunable parameter is different from second value of the tunable parameter.

5 . The method of claim 1 , wherein the trainable activation function comprises a tunable scaling parameter that scales an input of the trainable activation function, and training the neural network further comprises:

determining a value of the tunable scaling parameter based on the plurality of training samples.

6 . The method of claim 5 , wherein determining the value of the tunable scaling parameter comprises:

confining a domain of the tunable scaling parameter to a range from 1 to 10; and

determining the value of the tunable scaling parameter within the domain of the tunable scaling parameter.

7 . The method of claim 5 , wherein adding the trainable activation function with the differentiation operator to the one or more hidden layers of the plurality of hidden layers comprises:

adding the trainable activation function with the differentiation operator to a first hidden layer and a second hidden layer,

wherein training the neural network comprises:

 determining a first value of the tunable parameter and a second value of the tunable scaling parameter for the first hidden layer based on the plurality of training samples, and

 determining a third value of the tunable parameter and a fourth value of the tunable scaling parameter for the second hidden layer based on the plurality of training samples.

8 . One or more non-transitory computer-readable media storing instructions executable to perform operations for deep learning based on a trainable activation function, the operations comprising:

accessing a neural network comprising an input layer, a plurality of hidden layers, and an output layer;

adding the trainable activation function with a differentiation operator to one or more hidden layers of the plurality of hidden layers, the trainable activation function comprising a tunable parameter, the tunable parameter representing a fractional derivative order of the differentiation operator in the trainable activation function; and

training the neural network by inputting a plurality of training samples into the neural network, wherein training the neural network comprises:

determining a value of the tunable parameter based on the plurality of training samples.

9 . The one or more non-transitory computer-readable media of claim 8 , wherein determining the value of the tunable parameter based on the plurality of training samples comprises:

confining a domain of the tunable parameter to a range from 0 to 2; and

determining the value of the tunable parameter within the domain based on the plurality of training samples.

10 . The one or more non-transitory computer-readable media of claim 8 , wherein adding the trainable activation function with the differentiation operator to the one or more hidden layers of the plurality of hidden layers comprises:

adding the trainable activation function with the differentiation operator to a first hidden layer and a second hidden layer,

wherein training the neural network comprises:

determining a first value of the tunable parameter for the first hidden layer based on the plurality of training samples, and

determining a second value of the tunable parameter for the second hidden layer based on the plurality of training samples.

11 . The one or more non-transitory computer-readable media of claim 10 , wherein the first value of the tunable parameter is different from second value of the tunable parameter.

12 . The one or more non-transitory computer-readable media of claim 8 , wherein the trainable activation function comprises a tunable scaling parameter that scales an input of the trainable activation function, and training the neural network further comprises:

determining a value of the tunable scaling parameter based on the plurality of training samples.

13 . The one or more non-transitory computer-readable media of claim 12 , wherein determining the value of the tunable scaling parameter comprises:

confining a domain of the tunable scaling parameter to a range from 1 to 10; and

determining the value of the tunable scaling parameter within the domain of the tunable scaling parameter.

14 . The one or more non-transitory computer-readable media of claim 12 , wherein adding the trainable activation function with the differentiation operator to the one or more hidden layers of the plurality of hidden layers comprises:

adding the trainable activation function with the differentiation operator to a first hidden layer and a second hidden layer,

wherein training the neural network comprises:

 determining a first value of the tunable parameter and a second value of the tunable scaling parameter for the first hidden layer based on the plurality of training samples, and

 determining a third value of the tunable parameter and a fourth value of the tunable scaling parameter for the second hidden layer based on the plurality of training samples.

15 . An apparatus for deep learning based on a trainable activation function, the apparatus comprising:

a computer processor for executing computer program instructions; and

a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising:

 accessing a neural network comprising an input layer, a plurality of hidden layers, and an output layer,

 adding the trainable activation function with a differentiation operator to one or more hidden layers of the plurality of hidden layers, the trainable activation function comprising a tunable parameter, the tunable parameter representing a fractional derivative order of the differentiation operator in the trainable activation function, and

 training the neural network by inputting a plurality of training samples into the neural network, wherein training the neural network comprises:

determining a value of the tunable parameter based on the plurality of training samples.

16 . The apparatus of claim 15 , wherein determining the value of the tunable parameter based on the plurality of training samples comprises:

confining a domain of the tunable parameter to a range from 0 to 2; and

determining the value of the tunable parameter within the domain based on the plurality of training samples.

17 . The apparatus of claim 15 , wherein adding the trainable activation function with the differentiation operator to the one or more hidden layers of the plurality of hidden layers comprises:

adding the trainable activation function with the differentiation operator to a first hidden layer and a second hidden layer,

wherein training the neural network comprises:

determining a first value of the tunable parameter for the first hidden layer based on the plurality of training samples, and

determining a second value of the tunable parameter for the second hidden layer based on the plurality of training samples.

18 . The apparatus of claim 17 , wherein the first value of the tunable parameter is different from second value of the tunable parameter.

19 . The apparatus of claim 15 , wherein the trainable activation function comprises a tunable scaling parameter that scales an input of the trainable activation function, and training the neural network further comprises:

determining a value of the tunable scaling parameter based on the plurality of training samples.

20 . The apparatus of claim 19 , wherein determining the value of the tunable scaling parameter comprises:

confining a domain of the tunable scaling parameter to a range from 1 to 10; and

determining the value of the tunable scaling parameter within the domain of the tunable scaling parameter.