IP Library Granted Patent US 10,719,760
Granted Patent B2
US 10,719,760 · App. 15/482,793 · Granted Jul 21, 2020

Neural network scheduling mechanism

Inventors: Liwei Ma (Beijing, CN); Nadathur Rajagopalan Satish (Santa Clara, CA); Jeremy Bottleson (Rancho Cordova, CA); Farshad Akhbari (Chandler, AZ); Eriko Nurvitadhi (Hillsboro, OR); Chandrasekaran Sakthivel (Sunnyvale, CA); Barath Lakshmanan (Chandler, AZ); Jingyi Jin (Folsom, CA); Justin E. Gottschlich (Santa Clara, CA); Michael Strickland (Sunnyvale, CA)
Assignee: Intel Corporation
G06N3/0445G06F9/5038G06N3/0454G06N3/063G06N3/084G06F2209/5021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,719,760
App. No.
15/482,793
Filed
Apr 9, 2017
Granted
Jul 21, 2020
Kind
B2
Art Unit
2123
USPC
706/25
Abstract

An apparatus to facilitate workload scheduling is disclosed. The apparatus includes one or more clients, one or more processing units to processes workloads received from the one or more clients, including hardware resources and scheduling logic to schedule direct access of the hardware resources to the one or more clients to process the workloads.

Claims (26)

1. An apparatus to facilitate workload scheduling comprising:

one or more clients; and

one or more general purpose graphics processing units to processes general purpose graphics workloads received from the one or more clients, including:

hardware resources;

a scheduler to schedule direct access to the hardware resources to the one or more clients to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

a gather unit including a relative address table, the relative address table to store a map to non-zero data values associated with the one or more clients, wherein the gather unit is to gather the non-zero data values and bypass zero data values.

2. The apparatus of claim 1 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

3. The apparatus of claim 2 , wherein the scheduler schedules access to each of the one or more clients via a Kernel Mode Driver (KMD).

4. The apparatus of claim 2 , wherein the scheduler provides access to each of the one or more clients to the hardware resources based on priority and a submission client type.

5. The apparatus of claim 2 , further comprising driver logic to access the one or more processing units, wherein each of the one or more clients are registered at the driver logic.

6. The apparatus of claim 5 , wherein each of the one or more clients receives a function pointer to enable direct access to the hardware resources.

7. The apparatus of claim 5 , wherein each of the one or more clients comprise an input interface to the one or more processing units.

8. A method to facilitate workload scheduling, comprising:

receiving requests from one or more clients to access hardware resources of a general purpose graphics processing unit;

scheduling direct access of the hardware resources to a first client of the one or more clients to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the one or more clients while bypassing zero data values associated with the one or more clients, the gathering performed via a relative address table, the relative address table to store a map to the non-zero data values.

9. The method of claim 8 , wherein access to the first client is scheduled via a Kernel Mode Driver (KMD).

10. The method of claim 8 , wherein access is provided to the first client based on a priority and a submission client type.

11. The method of claim 10 , further comprising registering the first client with driver logic associated with the processing unit.

12. At least one non-transitory computer readable medium having instructions, which when executed by one or more processors, the one or more processors including a general purpose graphics processing unit, cause the one or more processors to:

receive requests from one or more clients to access hardware resources of the general purpose graphics processing unit to process workloads associated with the one or more clients;

schedule direct access of the hardware resources to a first client of the one or more clients to process the workloads, wherein the one or more clients are each associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the one or more clients while bypassing zero data values associated with the one or more clients, the gathering performed via a relative address table, the relative address table to store a map to the non-zero data values.

13. The computer readable medium of claim 12 , wherein access to the first client is scheduled via a Kernel Mode Driver (KMD).

14. The computer readable medium of claim 12 , wherein access is provided to the first client based on a priority and a submission client type.

15. The computer readable medium of claim 14 , having instructions, which when executed by the one or more processors, further causes the processors to register the first client with driver logic associated with the processing unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2017
From: MA, LIWEI; SATISH, NADATHUR RAJAGOPALAN; BOTTLESON, JEREMY; AKHBARI, FARSHAD; NURVITADHI, ERIKO; SAKTHIVEL, CHANDRASEKARAN; LAKSHMANAN, BARATH; JIN, JINGYI; GOTTSCHLICH, JUSTIN; STRICKLAND, MICHAEL
To: INTEL CORPORATION
Reel/Frame 042812/0102 →
Continuity (1)
Related Publication 20180293490A1 · Oct 11, 2018
Cited By (1)
US 12,210,959