IP Library Granted Patent US 11,809,978
Granted Patent B2
US 11,809,978 · App. 17/723,074 · Granted Nov 7, 2023

Neural network scheduling mechanism

Inventors: Liwei Ma (Beijing, CN); Nadathur Rajagopalan Satish (Santa Clara, CA); Jeremy Bottleson (Rancho Cordova, CA); Farshad Akhbari (Chandler, AZ); Eriko Nurvitadhi (Hillsboro, OR); Chandrasekaran Sakthivel (Sunnyvale, CA); Barath Lakshmanan (Chandler, AZ); Jingyi Jin (Folsom, CA); Justin E. Gottschlich (Santa Clara, CA); Michael Strickland (Sunnyvale, CA)
Assignee: Intel Corporation
G06N3/044G06F9/5038G06N3/045G06N3/063G06N3/084G06F2209/5021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,809,978
App. No.
17/723,074
Granted
Nov 7, 2023
Kind
B2
Abstract

An apparatus to facilitate workload scheduling is disclosed. The apparatus includes one or more clients, one or more processing units to processes workloads received from the one or more clients, including hardware resources and scheduling logic to schedule direct access of the hardware resources to the one or more clients to process the workloads.

Claims (33)

1. An apparatus to facilitate workload scheduling comprising:

an interconnect fabric; and

one or more general purpose graphics processing units coupled with the interconnect fabric, the one or more general purpose graphics processing units to processes general purpose graphics workloads received from one or more clients via the interconnect fabric, the one or more general purpose graphics processing units including:

hardware resources;

a scheduler to schedule direct access to the hardware resources to process the workloads on behalf of the one or more clients, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

a gather unit to bypass zero data values and gather non-zero data values associated with the one or more clients, the non-zero data values stored sparsely in memory.

2. The apparatus of claim 1 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

3. The apparatus of claim 2 , wherein convolutional kernel is an irregular convolutional kernel and the hardware resources are configured to multiply the data values for the convolutional kernel by the data elements of the feature map.

4. The apparatus of claim 2 , wherein the scheduler provides access to each of the one or more clients to the hardware resources based on priority and a submission client type.

5. The apparatus of claim 2 , wherein the one or more general purpose graphics processing units are associated with driver logic to facilitate access to the one or more general purpose graphics processing units and each of the one or more clients are registered to the driver logic.

6. The apparatus of claim 5 , wherein each of the one or more clients receives a function pointer to enable direct access to the hardware resources.

7. The apparatus of claim 5 , wherein each of the one or more clients comprise an input interface to the one or more general purpose graphics processing units.

8. The apparatus of claim 1 , wherein the gather unit is to store a map to the non-zero data values associated with the one or more clients and gather the non-zero data values based on the map.

9. A method to facilitate workload scheduling, comprising:

receiving a request to access hardware resources of a general purpose graphics processing unit;

scheduling direct access to the hardware resources to enable a client to process a workload provided by the client, wherein the client is associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the client while bypassing zero data values associated with the client, the non-zero data values stored sparsely in memory.

10. The method of claim 9 , wherein the gathering is performed via a map to the non-zero data values.

11. The method of claim 9 , further comprising scheduling direct access to the hardware resources via a Kernel Mode Driver (KMD) associated with the general purpose graphics processing unit.

12. The method of claim 9 , wherein access is provided to the client based on a priority and a submission client type.

13. The method of claim 12 , further comprising registering the client with driver logic associated with the general purpose graphics processing unit.

14. The method as in claim 9 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

15. The method as in claim 14 , further comprising multiplying, via the hardware resources of the general purpose graphics processing unit, the data values for the convolutional kernel by the data elements of the feature map.

16. A data processing system comprising:

memory to store instructions;

one or more processors configured to execute the instructions, wherein the one or more processors include a graphics processor and the instructions configure the one or more processors to:

receive a request to access hardware resources of a general purpose graphics processing unit;

schedule direct access to the hardware resources to enable a client to process a workload provided by the client, wherein the client is associated with a precompiled neural network (NN) kernel; and

gather, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the client while bypassing zero data values associated with the client, the non-zero data values stored sparsely in memory.

17. The data processing system as in claim 16 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

18. The data processing system of claim 17 , wherein convolutional kernel is an irregular convolutional kernel and the hardware resources are configured to multiply the data values for the convolutional kernel by the data elements of the feature map.

19. The data processing system of claim 17 , wherein the scheduler provides access to each of the one or more clients to the hardware resources based on priority and a submission client type.

20. The data processing system of claim 17 , wherein the one or more general purpose graphics processing units are associated with driver logic to facilitate access to the one or more general purpose graphics processing units and each of the one or more clients are registered to the driver logic.

Continuity (3)
Continuation 16918220 · Jul 1, 2020
Continuation 15482793 · Apr 9, 2017
Related Publication 20220327357A1 · Oct 13, 2022