IP Library Granted Patent US 11,880,715
Granted Patent B2
US 11,880,715 · App. 17/222,543 · Granted Jan 23, 2024

Method and system for opportunistic load balancing in neural networks using metadata

Inventors: Nicholas Malaya (Austin, TX); Yasuko Eckert (Bellevue, WA)
Assignee: Advanced Micro Devices, Inc.
G06F9/5044G06F9/505G06F9/5066G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,880,715
App. No.
17/222,543
Granted
Jan 23, 2024
Kind
B2
Abstract

Methods and systems for load balancing in a neural network system using metadata are disclosed. Any one or a combination of one or more kernels, one or more neurons, and one or more layers of the neural network system are tagged with metadata. A scheduler detects whether there are neurons that are available to execute. The scheduler uses the metadata to schedule and load balance computations across compute resources and available resources.

Claims (31)

1. A method for load balancing in a neural network system that improves computational efficiency of the neural network system, the method comprising:

tagging any one or a combination of one or more kernels, one or more neurons, and one or more layers of the neural network system with metadata;

detecting that there are resources available for execution of a neuron in a subsequent layer of the neural network system; and

in response to the detecting that there are resources available for execution:

performing the load balancing in the neural network system using the metadata, and

executing computations of the neuron in the subsequent layer of the neural network system according to the load balancing.

2. The method of claim 1 , wherein the load balancing includes the subsequent layer and one or more layers subsequent to the subsequent layer.

3. The method of claim 1 , wherein the load balancing is performed while a current layer is being processed.

4. The method of claim 1 , wherein the metadata indicates any one or a combination of a kernel size, a filter size, a dropout layer, a number of neurons present in a layer, neuron readiness, or an activation function.

5. The method of claim 4 , wherein the neuron readiness accounts for dependencies between neurons.

6. The method of claim 5 , wherein the neuron readiness is updated after a neuron associated with the neuron readiness is pruned.

7. The method of claim 1 , wherein the metadata is tagged by any one or a combination of an application, a framework, or a user.

8. The method of claim 1 , wherein the metadata is stored in any one or a combination of an instruction, a scalar register, a memory, and a hardware table.

9. The method of claim 1 , further comprising:

obtaining representative computational costs for the neural network system, wherein a computational cost is determined from the representative computational costs; and

storing the computational costs for the neural network system as a component of the metadata.

10. A system for load balancing in a neural network system that improves computational efficiency of the neural network system, comprising:

one or more kernels, one or more neurons, and one or more layers, wherein any one or a combination of the one or more kernels, neurons, and layers of the neural network system are tagged with metadata; and

a scheduler connected to the neural network system, where the scheduler is configured to:

detect that there are resources available for execution of a neuron in a subsequent layer of the neural network system; and

in response to detecting that there are resources available for execution:

perform the load balancing of the neural network system using the metadata, and

execute computations of the neuron in the subsequent layer of the neural network system according to the load balancing.

11. The system of claim 10 , wherein the load balancing includes the subsequent layer and one or more layers subsequent to the subsequent layer.

12. The system of claim 10 , wherein the load balancing is performed while a current layer is being processed.

13. The system of claim 10 , wherein the metadata indicates any one or a combination of a kernel size, a filter size, a dropout layer, a number of neurons present in a layer, neuron readiness, or an activation function.

14. The system of claim 13 , wherein the neuron readiness accounts for dependencies between neurons.

15. The system of claim 14 , wherein the neuron readiness is updated after a neuron associated with the neuron readiness is pruned.

16. The system of claim 10 , wherein the metadata is tagged by any one or a combination of an application, a framework, or a user.

17. The system of claim 10 , wherein the metadata is stored in at least an instruction, a scalar register, a memory, and a hardware table.

18. The system of claim 10 , wherein the metadata includes computational costs for the neural network system, a computational cost being determined from obtained representative computational costs for the neural network system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2024
From: ADVANCED MICRO DEVICES, INC.
To: ONESTA IP, LLC
Reel/Frame 069381/0951 →
Continuity (2)
Continuation 16019374 · Jun 26, 2018
Related Publication 20210224130A1 · Jul 22, 2021