IP Library Granted Patent US 11,315,007
Granted Patent B2
US 11,315,007 · App. 16/918,220 · Granted Apr 26, 2022

Neural network scheduling mechanism

Inventors: Liwei Ma (Beijing, CN); Nadathur Rajagopalan Satish (Santa Clara, CA); Jeremy Bottleson (Rancho Cordova, CA); Farshad Akhbari (Chandler, AZ); Eriko Nurvitadhi (Hillsboro, OR); Chandrasekaran Sakthivel (Sunnyvale, CA); Barath Lakshmanan (Chandler, AZ); Jingyi Jin (Folsom, CA); Justin E. Gottschlich (Santa Clara, CA); Michael Strikland (Sunnyvale, CA)
Assignee: Intel Corporation
G06N3/0445G06F9/5038G06N3/0454G06N3/063G06N3/084G06F2209/5021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,315,007
App. No.
16/918,220
Filed
Jul 1, 2020
Granted
Apr 26, 2022
Kind
B2
Art Unit
2849
USPC
706/25
Abstract

An apparatus to facilitate workload scheduling is disclosed. The apparatus includes one or more clients, one or more processing units to processes workloads received from the one or more clients, including hardware resources and scheduling logic to schedule direct access of the hardware resources to the one or more clients to process the workloads.

Claims (35)

1. An apparatus to facilitate workload scheduling comprising:

one or more clients; and

one or more general purpose graphics processing units to processes general purpose graphics workloads received from the one or more clients, the one or more general purpose graphics processing units including:

hardware resources;

a scheduler to schedule, to the one or more clients, direct access to the hardware resources to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

a gather unit to bypass zero data values and gather non-zero data values associated with the one or more clients, the non-zero data values stored sparsely in memory.

2. The apparatus of claim 1 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

3. The apparatus of claim 2 , wherein the scheduler schedules access to each of the one or more clients in association with a Kernel Mode Driver (KMD).

4. The apparatus of claim 2 , wherein the scheduler provides access to each of the one or more clients to the hardware resources based on priority and a submission client type.

5. The apparatus of claim 2 , wherein the one or more general purpose graphics processing units are associated with driver logic to facilitate access to the one or more general purpose graphics processing units and each of the one or more clients are registered to the driver logic.

6. The apparatus of claim 5 , wherein each of the one or more clients receives a function pointer to enable direct access to the hardware resources.

7. The apparatus of claim 5 , wherein each of the one or more clients comprise an input interface to the one or more general purpose graphics processing units.

8. The apparatus of claim 1 , wherein the gather unit is to store a map to the non-zero data values associated with the one or more clients and gather the non-zero data values based on the map.

9. A method to facilitate workload scheduling, comprising:

receiving requests from one or more clients to access hardware resources of a general purpose graphics processing unit;

scheduling direct access of the hardware resources to a first client of the one or more clients to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the one or more clients while bypassing zero data values associated with the one or more clients, the non-zero data values stored sparsely in memory.

10. The method of claim 9 , wherein the gathering is performed via a map to the non-zero data values.

11. The method of claim 9 , wherein access to the first client is scheduled via a Kernel Mode Driver (KMD).

12. The method of claim 9 , wherein access is provided to the first client based on a priority and a submission client type.

13. The method of claim 12 , further comprising registering the first client with driver logic associated with the processing unit.

14. At least one non-transitory computer readable medium having instructions, which when executed by one or more processors, the one or more processors including a general purpose graphics processing unit, cause the one or more processors to:

receive requests from one or more clients to access hardware resources of the general purpose graphics processing unit to process workloads associated with the one or more clients;

schedule direct access of the hardware resources to a first client of the one or more clients to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the one or more clients while bypassing zero data values associated with the one or more clients, the non-zero data values stored sparsely in memory.

15. The computer readable medium of claim 14 , wherein the gathering is performed via a map to the non-zero data values.

16. The computer readable medium of claim 14 , wherein access to the first client is scheduled via a Kernel Mode Driver (KMD).

17. The computer readable medium of claim 14 , wherein access is provided to the first client based on a priority and a submission client type.

18. The computer readable medium of claim 17 , having instructions, which when executed by the one or more processors, further causes the processors to register the first client with driver logic associated with the processing unit.

19. A data processing system comprising:

one or more clients; and

one or more general purpose graphics processing units to processes general purpose graphics workloads received from the one or more clients, the one or more general purpose graphics processing units including:

hardware resources;

a scheduler to schedule, to the one or more clients, direct access to the hardware resources to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel and each of the one or more clients receives a function pointer to enable direct access to the hardware resources; and

a gather unit to bypass zero data values and gather non-zero data values associated with the one or more clients, wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

Continuity (2)
Continuation 15482793 · Apr 9, 2017
Related Publication 20200394498A1 · Dec 17, 2020
Cited By (1)
US 12,499,347