IP Library Granted Patent US 10,417,731
Granted Patent B2
US 10,417,731 · App. 15/494,886 · Granted Sep 17, 2019

Compute optimization mechanism for deep neural networks

Inventors: Prasoonkumar Surti (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Joydeep Ray (Folsom, CA); Ben J. Ashbaugh (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Rajkishore Barik (Santa Clara, CA); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Altug Koker (El Dorado Hills, CA); Nadathur Rajagopalan Satish (Santa Clara, CA); Farshad Akhbari (Chandler, AZ); Dukhwan Kim (San Jose, CA); Wenyin Fu (Folsom, CA); Travis T. Schluessler (Hillsboro, OR); Josh B. Mastronarde (Sacramento, CA); Linda L. Hurd (Cool, CA); John H. Feit (Folsom, CA); Jeffery S. Boles (Folsom, CA); Adam T. Lake (Portland, OR); Karthik Vaidyanathan (Berkeley, CA); Devan Burke (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Abhishek R. Appu (El Dorado Hills, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F9/45533G06F9/5061G06F9/5094G06N3/0445G06N3/0454G06N3/063G06N3/084G06F8/41G06F2009/45583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,417,731
App. No.
15/494,886
Filed
Apr 24, 2017
Granted
Sep 17, 2019
Kind
B2
Art Unit
2618
USPC
345/520
Abstract

An apparatus to facilitate compute optimization is disclosed. The apparatus includes a plurality of processing units each comprising a plurality of execution units (EUs), wherein the plurality of EUs comprise a first EU type and a second EU type.

Claims (46)

1. An apparatus, comprising:

a plurality of processing units each comprising a plurality of execution units (EUs), each EU of the plurality of EUs to execute one or more threads, wherein the plurality of EUs comprise a first EU type and a second EU type, the second EU type being different from the first EU type, wherein:

each EU of the first EU type is capable of processing a first number of threads and supporting a first number of registers;

each EU of the second EU type is capable of processing a second number of threads and supporting a second number of registers; and

the first number of threads is greater than the second number of threads, and the second number of registers is greater than the first number of registers.

2. The apparatus of claim 1 , wherein the plurality of processing units comprise:

a first processing unit including a plurality of EUs of the first type; and

a second processing unit including a plurality of EUs of the second type.

3. The apparatus of claim 1 , wherein the plurality of processing units comprise:

a first processing unit including:

a first set of EUs of the first type; and

a second set of EUs of the second type; and

a second processing unit including:

a third set of EUs of the first type; and

a fourth set of EUs of the second type.

4. The apparatus of claim 1 , wherein the apparatus is to select one or more EUs of the plurality of EUs to be implemented to execute a workload.

5. The apparatus of claim 4 , wherein the apparatus is to select one or more EUs of the first type to process a first type of application workload and selects one or more EUs of the second type to process a second type of application workload.

6. The apparatus of claim 5 , wherein the first type of application workload includes a three-dimensional (3D) application and the second type of application workload includes a media application.

7. The apparatus of claim 5 , wherein the selections of types of EUs are to remain static during a lifetime of respective applications in the first type of application workload and in the second type of application workload.

8. The apparatus of claim 1 , further comprising a memory device, wherein the plurality of processing units are included in the memory device.

9. The apparatus of claim 8 , wherein the memory device comprises a high bandwidth memory (HBM).

10. The apparatus of claim 9 , wherein the HBM comprises:

a first memory channel; and

a first processing unit of the plurality of processing units included in the first memory channel.

11. The apparatus of claim 1 , further comprising a register file implemented to facilitate matrix-vector transformations performed by one or more of the plurality of processing units.

12. The apparatus of claim 1 , further comprising a shared local memory (SLM) implemented to facilitate matrix-vector transformations performed by one or more of the plurality of processing units.

13. A graphics processor comprising:

a plurality of processing units each comprising a plurality of execution units (EUs), each EU of the plurality of EUs to execute one or more threads, wherein the plurality of EUs comprise a first EU type and a second EU type, the second EU type being different from the first EU type, wherein:

each EU of the first EU type is capable of processing a first number of threads and supporting a first number of registers,

each EU of the second EU type is capable of processing a second number of threads and supporting a second number of registers, and

the first number of threads is greater than the second number of threads, and the second number of registers is greater than the first number of registers;

a first processing unit including a first set of execution units (EUs); and

a second processing unit including a second set of EUs, wherein the first and second sets of EUs are comprised of the first EU type and the second EU type.

14. The graphics processor of claim 13 , wherein the first set of EUs comprise a plurality of EUs of the first type and the second set of EUs comprise a plurality of EUs of the second type.

15. The graphics processor of claim 13 , wherein the first and second sets of EUs each comprise one or more EUs of the first type and one or more EUs of the second type.

16. The graphics processor of claim 13 , wherein the graphics processor is to select one or more EUs of the plurality of EUs to be implemented to execute a workload.

17. The graphics processor of claim 16 , wherein the graphics processor is to select one or more EUs of the first type to process a first type of application workload and selects one or more EUs of the second type to process a second type of application workload.

18. The graphics processor of claim 17 , wherein the first type of application workload includes a three-dimensional (3D) application workload and the second type of application workload includes a media application workload.

19. The graphics processor of claim 17 , wherein the selections of types of EUs are to remain static during a lifetime of respective applications in the first type of application workload and in the second type of application workload.

20. The graphics processor of claim 13 , further comprising a memory device, wherein the plurality of processing units are included in the memory device.

21. The graphics processor of claim 20 , wherein the memory device comprises a high bandwidth memory (HBM).

22. The graphics processor of claim 21 , wherein the HBM comprises:

a first memory channel; and

a first processing unit of the plurality of processing units included in the first memory channel.

23. The graphics processor of claim 13 , further comprising a register file implemented to facilitate matrix-vector transformations performed by one or more of the plurality of processing units.

24. The graphics processor of claim 13 , further comprising a shared local memory (SLM) implemented to facilitate matrix-vector transformations performed by one or more of the plurality of processing units.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2017
From: SURTI, PRASOONKUMAR; SRINIVASA, NARAYAN; CHEN, FENG; RAY, JOYDEEP; ASHBAUGH, BEN J.; GALOPPO VON BORRIES, NICOLAS C.; NURVITADHI, ERIKO; VEMBU, BALAJI; LIN, TSUNG-HAN; SINHA, KAMAL; BARIK, RAJKISHORE; BAGHSORKHI, SARA S.; GOTTSCHLICH, JUSTIN E.; KOKER, ALTUG; SATISH, NADATHUR RAJAGOPALAN; AKHBARI, FARSHAD; KIM, DUKHWAN; FU, WENYIN; SCHLUESSLER, TRAVIS T.; MASTRONARDE, JOSH B.; HURD, LINDA L.; FEIT, JOHN H.; BOLES, JEFFREY S.; LAKE, ADAM T.; APPU, ABHISHEK R.; VAIDYANATHAN, KARTHIK; BURKE, DEVAN; MAIYURAN, SUBRAMANIAM
To: INTEL CORPORATION
Reel/Frame 044339/0942 →
Continuity (1)
Related Publication 20180308200A1 · Oct 25, 2018
Cited By (2)
US 12,198,221 US 12,580,977