IP Library › Granted Patent US 11,972,354
Granted Patent B2
US 11,972,354 · App. 17/672,543 · Granted Apr 30, 2024

Representing a neural network utilizing paths within the network to improve a performance of the neural network

Inventors: Alexander Keller (Berlin, DE); Gonçalo Filipe Torcato Mordido (Potsdam, DE); Noah Jonathan Gamboa (Menlo Park, CA); Matthijs Jules Van Keirsbilck (Berlin, DE)
Assignee: NVIDIA CORPORATION
G06N3/088G06N3/082G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,972,354
App. No.
17/672,543
Granted
Apr 30, 2024
Kind
B2
Abstract

Artificial neural networks (ANNs) are computing systems that imitate a human brain by learning to perform tasks by considering examples. By representing an artificial neural network utilizing individual paths each connecting an input of the ANN to an output of the ANN, a complexity of the ANN may be reduced, and the ANN may be trained and implemented in a much faster manner when compared to an implementation using fully connected ANN graphs.

Claims (47)

1. A method comprising, at a device:

creating an artificial neural network (ANN) by sampling a plurality of paths from another ANN, wherein the sampling is proportional to at least one of:

a plurality of given weights of layers of the ANN, or

a plurality of given activations of layers of the ANN; and

normalizing the ANN by:

normalizing a plurality of weights of a first layer of the ANN with a linear factor; and

propagating the normalized plurality of weights to successive layers of the ANN by multiplying weights of each of the successive layers by the linear factor of a previous layer;

wherein a last layer of the ANN stores resulting weights to scale outputs of the ANN.

2. The method of claim 1 , wherein the ANN includes a ReLU activation function.

3. The method of claim 1 , wherein the ANN includes a leaky ReLU activation function.

4. The method of claim 1 , wherein the ANN includes a maxpool activation function.

5. The method of claim 1 , wherein the ANN includes an absolute value activation function.

6. The method of claim 1 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given weights of layers of the ANN.

7. The method of claim 1 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given activations of layers of the ANN.

8. A system comprising:

a hardware processor of a device that is configured to:

create an artificial neural network (ANN) by sampling a plurality of paths from another ANN, wherein the sampling is proportional to at least one of:

a plurality of given weights of layers of the ANN, or

a plurality of given activations of layers of the ANN; and

normalize the ANN by:

normalizing a plurality of weights of a first layer of the ANN with a linear factor; and

propagating the normalized plurality of weights to successive layers of the ANN by multiplying weights of each of the successive layers by the linear factor of a previous layer;

wherein a last layer of the ANN stores resulting weights to scale outputs of the ANN.

9. The system of claim 8 , wherein the ANN includes a ReLU activation function.

10. The system of claim 8 , wherein the ANN includes a leaky ReLU activation function.

11. The system of claim 8 , wherein the ANN includes a maxpool activation function.

12. The system of claim 8 , wherein the ANN includes an absolute value activation function.

13. The system of claim 8 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given weights of layers of the ANN.

14. The system of claim 8 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given activations of layers of the ANN.

15. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of a device, causes the processor to:

create an artificial neural network (ANN) by sampling a plurality of paths from another ANN, wherein the sampling is proportional to at least one of:

a plurality of given weights of layers of the ANN, or

a plurality of given activations of layers of the ANN; and

normalize the ANN by:

normalizing a plurality of weights of a first layer of the ANN with a linear factor; and

propagating the normalized plurality of weights to successive layers of the ANN by multiplying weights of each of the successive layers by the linear factor of a previous layer;

wherein a last layer of the ANN stores resulting weights to scale outputs of the ANN.

16. The computer-readable storage medium of claim 15 , wherein the ANN includes a ReLU activation function.

17. The computer-readable storage medium of claim 15 , wherein the ANN includes a leaky ReLU activation function.

18. The computer-readable storage medium of claim 15 , wherein the ANN includes a maxpool activation function.

19. The computer-readable storage medium of claim 15 , wherein the ANN includes an absolute value activation function.

20. The computer-readable storage medium of claim 15 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given weights of layers of the ANN.

21. The computer-readable storage medium of claim 15 , wherein the ANN is created by sampling the plurality of paths from the another ANN, wherein the sampling is proportional to the plurality of given activations of layers of the ANN.

22. The method of claim 1 , wherein the ANN is an approximation operator, and wherein the ANN is normalized to control the operator norm.

23. The method of claim 22 , wherein the operator norm is a Euclidian norm.

24. The method of claim 23 , wherein the Euclidian norm is used to separate a length of a weight vector from its direction in order to accelerate training.

25. The method of claim 1 , wherein the linear function is used to normalize the weights such that the normalized weights form a discrete probability distribution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2022
From: KELLER, ALEXANDER; TORCATO MORDIDO, GONÇALO FILIPE; GAMBOA, NOAH JONATHAN; VAN KEIRSBILCK, MATTHIJS JULES
To: NVIDIA CORPORATION
Reel/Frame 059785/0761 →
Continuity (3)
Division 16352596 · Mar 13, 2019
Provisional Application 62648263 · Mar 26, 2018
Related Publication 20220172072A1 · Jun 2, 2022