IP Library Granted Patent US 12,499,347
Granted Patent B2
US 12,499,347 · App. 18/471,843 · Granted Dec 16, 2025

Neural network scheduling mechanism

Inventors: Liwei Ma (Beijing, CN); Nadathur Rajagopalan Satish (Santa Clara, CA); Jeremy Bottleson (Rancho Cordova, CA); Farshad Akhbari (Chandler, AZ); Eriko Nurvitadhi (Hillsboro, OR); Chandrasekaran Sakthivel (Sunnyvale, CA); Barath Lakshmanan (Chandler, AZ); Jingyi Jin (Folsom, CA); Justin E. Gottschlich (Santa Clara, CA); Michael Strickland (Sunnyvale, CA)
Assignee: Intel Corporation
G06N3/044G06F9/5038G06N3/045G06N3/063G06N3/084G06F2209/5021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,347
App. No.
18/471,843
Granted
Dec 16, 2025
Kind
B2
Abstract

An apparatus to facilitate workload scheduling is disclosed. The apparatus includes one or more clients, one or more processing units to processes workloads received from the one or more clients, including hardware resources and scheduling logic to schedule direct access of the hardware resources to the one or more clients to process the workloads.

Claims (34)

1 . A graphics processing unit comprising:

a shared memory;

a memory interface coupled with the shared memory; and

a processing cluster including a plurality of graphics multiprocessors internal to the graphics processing unit, the processing cluster coupled with the shared memory via the memory interface, the plurality of graphics multiprocessors coupled via a data interconnect, the data interconnect to facilitate exchange of data between the plurality of graphics multiprocessors during cooperative execution of a workload, the plurality of graphics multiprocessors configured to process workloads received for execution, wherein a graphics multiprocessor of the plurality of graphics multiprocessors includes:

a plurality of processing engines;

a scheduler to schedule direct access to the plurality of processing engines to process the workloads, wherein the workloads are each associated with a precompiled neural network (NN) kernel; and

a gather unit to bypass zero data values and gather non-zero data values associated with the workloads, the non-zero data values stored sparsely in memory.

2 . The graphics processing unit of claim 1 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

3 . The graphics processing unit of claim 2 , wherein convolutional kernel is an irregular convolutional kernel and the plurality of processing engines are configured to multiply the data values for the convolutional kernel by the data elements of the feature map.

4 . The graphics processing unit of claim 2 , wherein the scheduler is to schedule the workloads to the plurality of processing engines based on priority and a submission type associated with the workloads.

5 . The graphics processing unit of claim 2 , wherein the plurality of graphics multiprocessors are associated with driver logic to facilitate access to the plurality of graphics multiprocessors by one or more clients, the one or more clients registered to the driver logic, the one or more clients to bypass an operating system to access the plurality of graphics multiprocessors via a function pointer received from the driver logic.

6 . The graphics processing unit of claim 5 , wherein each of the one or more clients receives a function pointer to enable direct access to the plurality of processing engines.

7 . The graphics processor processing unit of claim 5 , wherein each of the one or more clients include an input interface to the plurality of graphics multiprocessors.

8 . The graphics processing unit of claim 1 , wherein the gather unit is to store a map to the non-zero data values and gather the non-zero data values based on the map.

9 . A method to facilitate workload scheduling, comprising:

receiving a request to access plurality of processing engines of a general purpose graphics processing unit, the general purpose graphics processing unit including a processing cluster having a plurality of graphics multiprocessors internal to the general purpose graphics processing unit, the plurality of graphics multiprocessors coupled via a data interconnect, the data interconnect to facilitate exchange of data between the plurality of graphics multiprocessors during cooperative execution of a workload, and a graphics multiprocessor of the plurality of graphics multiprocessors includes the plurality of processing engines;

scheduling direct access to the plurality of processing engines to enable a client to process a workload provided by the client, wherein the workload is associated with a precompiled neural network (NN) kernel; and

gathering, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the client while bypassing zero data values associated with the client, the non-zero data values stored sparsely in memory.

10 . The method of claim 9 , further comprising gathering the non-zero data values via a map to the non-zero data values.

11 . The method of claim 9 , further comprising bypassing an operating system and scheduling direct access to the plurality of processing engines via a Kernel Mode Driver (KMD) associated with the general purpose graphics processing unit, wherein the KMD provides a function pointer to enable direct access to the plurality of processing engines.

12 . The method of claim 9 , wherein access is provided to the client based on a priority and a submission client type.

13 . The method of claim 12 , further comprising registering the client with driver logic associated with the general purpose graphics processing unit.

14 . The method as in claim 9 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

15 . The method as in claim 14 , further comprising multiplying, via the plurality of processing engines of the general purpose graphics processing unit, the data values for the convolutional kernel by the data elements of the feature map.

16 . A data processing system comprising:

memory to store instructions; and

one or more processors configured to execute the instructions, wherein the one or more processors include a graphics processing unit including a processing cluster having a plurality of graphics multiprocessors internal to the graphics processing unit, the plurality of graphics multiprocessors coupled via a data interconnect, the data interconnect to facilitate exchange of data between the plurality of graphics multiprocessors during cooperative execution of a workload, and the instructions configure the one or more processors to:

receive a request to access plurality of processing engines of a general purpose graphics processing unit;

schedule, via a scheduler, direct access to the plurality of processing engines to enable the graphics processing unit to process a workload provided by a client, wherein the workload is associated with a precompiled neural network (NN) kernel; and

gather, via a gather unit of the general purpose graphics processing unit, non-zero data values associated with the client while bypassing zero data values associated with the client, the non-zero data values stored sparsely in memory.

17 . The data processing system as in claim 16 , wherein the non-zero data values are data values for a convolutional kernel to be multiplied by data elements of a feature map.

18 . The data processing system of claim 17 , wherein convolutional kernel is an irregular convolutional kernel and the plurality of processing engines are configured to multiply the data values for the convolutional kernel by the data elements of the feature map.

19 . The data processing system of claim 17 , wherein the scheduler is to provide access to the plurality of processing engines based on a priority and a submission client type.

20 . The data processing system of claim 17 , wherein the graphics processing unit is associated with driver logic to facilitate access to the plurality of processing engines and the client is registered to the driver logic, the client to bypass an operating system to access the plurality of graphics multiprocessors via a function pointer received from the driver logic.

Continuity (4)
Continuation 17723074 · Apr 18, 2022
Continuation 16918220 · Jul 1, 2020
Continuation 15482793 · Apr 9, 2017
Related Publication 20240086683A1 · Mar 14, 2024
References Cited (41)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 10936569B1 · Baskaran · 2021 [cited by examiner]
US 11315007B2 · Ma · 2022 [cited by applicant]
US 20120075319A1 · Dally · 2012 [cited by examiner]
US 20140259016A1 · Lottes · 2014 [cited by examiner]
US 20150341218A1 · Chen · 2015 [cited by applicant]
US 20150379424A1 · Dirac et al. · 2015 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160148078A1 · Shen et al. · 2016 [cited by applicant]
US 20160179574A1 · Merrill, III · 2016 [cited by examiner]
US 20160180486A1 · Rao et al. · 2016 [cited by applicant]
US 20170039396A1 · Sharp · 2017 [cited by examiner]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180278588A1 · Cela · 2018 [cited by examiner]
CN 103383644A · 2013 [cited by applicant]
CN 108694689A · 2018 [cited by applicant]
EP 3385844A1 · 2018 [cited by applicant]
EP 3822788A1 · 2021 [cited by applicant]
Various authors, “Convolutional neural network”, Wikipedia, Mar. 14, 2016, all pages (Year: 2016). [cited by examiner]
Communication pursuant to Article 94(3) EPC for EP Application No. 18159604.0, mailed Aug. 27, 2019, 6 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC for EP Application No. 20214192.5, mailed May 27, 20122, 10 pages. [cited by applicant]
Extended European Search Report for European Application No. 18159604.0, mailed Sep. 3, 2018, 13 pages. [cited by applicant]
Extended European Search Report for European Application No. 20214192.5, mailed Apr. 1, 2021, 14 pages. [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/482,793 mailed Nov. 18, 2019, 19 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/918,220 mailed Sep. 15, 2021, 21 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/482,793 mailed Mar. 20, 2020, 8 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/918,220 mailed Jan. 5, 2022, 8 pages. [cited by applicant]
Office Action for European Patent Application No. 18159604.0, mailed Oct. 16, 2020, 11 pages. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/723,074 mailed Jul. 6, 2023, 8 pages. [cited by applicant]
CN Office Action for CN201810307374.6, Sep. 29, 2023, 9 pages. [cited by applicant]
Intention to grant for EP Application No. 20214192.5, 5 pages, Sep. 30, 2024. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) EPC for EP20214192, Mar. 7, 2024, 12 pages. [cited by applicant]
Notification of Publication for CN Application No. 201810307374.6 Jun. 10, 2025, 4 pages. [cited by applicant]