IP Library › Granted Patent US 11,461,637
Granted Patent B2
US 11,461,637 · App. 16/293,700 · Granted Oct 4, 2022

Real-time resource usage reduction in artificial neural networks

Inventors: Taro Sekiyama (Urayasu, JP); Kiyokuni Kawachiya (Yokohama, JP); Tung D. Le (Ichikawa, JP); Yasushi Negishi (Tokyo, JP)
Assignee: International Business Machines Corporation
G06N3/08G06N3/04G06N3/0454G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,637
App. No.
16/293,700
Granted
Oct 4, 2022
Kind
B2
Abstract

A generated algorithm used by a neural network is captured during execution of an iteration of the neural network. A candidate algorithm is identified based on the generated algorithm. A determination is made that the candidate algorithm utilizes less memory than the generated algorithm. Based on the determination the neural network is updated by replacing the generated algorithm with the candidate algorithm.

Claims (55)

1. A method comprising:

capturing, during execution of an iteration of a neural network, a generated algorithm used by the neural network during the iteration;

identifying, based on the generated algorithm, a candidate algorithm;

determining that the candidate algorithm utilizes less memory than the generated algorithm, wherein the determining includes testing a potential candidate algorithm with any memory that is not currently allocated to the neural network; and

updating, based on the determination, the neural network by replacing the generated algorithm of the neural network with the candidate algorithm.

2. The method of claim 1 , further comprising:

detecting, before the determining, that the neural network has finished execution of the iteration;

pausing execution of the neural network; and

resuming, after the updating, the neural network.

3. The method of claim 1 , wherein the execution of the iteration is a forward propagation and a backward propagation of the neural network.

4. The method of claim 3 , and wherein the method further comprises:

recording, during execution of the iteration, the memory utilization of the generated algorithm;

performing, outside of the neural network, a forward propagation and a backward propagation with the candidate algorithm; and

comparing the memory utilization of the recorded generated algorithm and the performed candidate algorithm.

5. The method of claim 1 , wherein the neural network is a define-by-run neural network and wherein the identifying the generated algorithm is an identifying an updated layout of the neural network.

6. The method of claim 1 , wherein the neural network is a convolutional neural network.

7. The method of claim 6 , wherein the generated algorithm is a convolutional algorithm.

8. The method of claim 1 , wherein the method further comprises:

monitoring memory usage allocated to the neural network after the replacement of the generated algorithm with the candidate algorithm; and

flagging, based on the monitoring the memory usage, a subset of the memory allocated to the neural network.

9. The method of claim 8 , wherein the neural network comprises a second algorithm, wherein the method further comprises:

assigning the flagged subset of the memory allocated to the neural network to the second algorithm of the neural network.

10. The method of claim 8 , wherein the method further comprises:

deallocating the flagged subset of memory.

11. A system comprising:

a memory; and

a processor, the processor communicatively coupled to the memory, the processor configured to perform a method comprising:

capturing, during execution of an iteration of a neural network, a first generated algorithm used by the neural network during the iteration, wherein the neural network includes a plurality of generated algorithms including the first generated algorithm;

identifying, based on the first generated algorithm, a candidate algorithm;

determining that the candidate algorithm utilizes less memory than the first generated algorithm; and

updating, based on the determination, the neural network by replacing the first generated algorithm of the neural network with the candidate algorithm.

12. The system of claim 11 , wherein the method further comprises:

detecting, before the determining, that the neural network has finished execution of the iteration;

pausing execution of the neural network; and

resuming, after the updating, the neural network.

13. The system of claim 11 , wherein the execution of the iteration is a forward propagation and a backward propagation of the neural network.

14. The system of claim 13 , and wherein the method further comprises:

recording, during execution of the iteration, the memory utilization of the first generated algorithm;

performing, outside of the neural network, a forward propagation and a backward propagation with the candidate algorithm; and

comparing the memory utilization of the recorded first generated algorithm and the performed candidate algorithm.

15. The system of claim 11 , wherein the neural network is a define-by-run neural network and wherein the identifying the generated algorithm is an identifying an updated layout of the neural network.

16. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to perform a method comprising:

capturing, during execution of an iteration of a neural network, a generated algorithm used by the neural network during the iteration, the neural network reside in a graphics card memory during execution;

identifying, based on the generated algorithm, a candidate algorithm;

determining that the candidate algorithm utilizes less graphics card memory than the generated algorithm; and

updating, based on the determination, the neural network by replacing the generated algorithm of the neural network with the candidate algorithm.

17. The computer program product of claim 16 , wherein the method further comprises:

detecting, before the determining, that the neural network has finished execution of the iteration;

pausing execution of the neural network; and

resuming, after the updating, the neural network.

18. The computer program product of claim 16 , wherein the neural network is a convolutional neural network.

19. The computer program product of claim 18 , wherein the generated algorithm is a convolutional algorithm.

20. The computer program product of claim 16 , wherein the method further comprises:

monitoring memory usage of the graphics card memory allocated to the neural network after the replacement of the generated algorithm with the candidate algorithm; and

flagging, based on the monitoring the memory usage, a subset of the graphics card memory allocated to the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2019
From: SEKIYAMA, TARO; KAWACHIYA, KIYOKUNI; LE, TUNG D.; NEGISHI, YASUSHI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048511/0927 →
Continuity (2)
Continuation 15622127 · Jun 14, 2017
Related Publication 20190205755A1 · Jul 4, 2019
Cited By (3)
US 12,353,987 US 12,602,576 US 12,737,603