IP Library Granted Patent US 10,417,734
Granted Patent B2
US 10,417,734 · App. 15/698,217 · Granted Sep 17, 2019

Compute optimization mechanism for deep neural networks

Inventors: Prasoonkumar Surti (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Joydeep Ray (Folsom, CA); Ben J. Ashbaugh (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsoom, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Rajkishore Barik (Santa Clara, CA); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Altug Koker (El Dorado Hills, CA); Nadathur Rajagopalan Satish (Santa Clara, CA); Farshad Akhbari (Chandler, AZ); Dukhwan Kim (San Jose, CA); Wenyin Fu (Folsom, CA); Travis T. Schluessler (Hillsboro, OR); Josh B. Mastronarde (Sacramento, CA); Linda L. Hurd (Cool, CA); John H. Feit (Folsom, CA); Jeffery S. Boles (Folsom, CA); Adam T. Lake (Portland, OR); Karthik Vaidyanathan (Berkeley, CA); Devan Burke (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Abhishek R. Appu (El Dorado Hills, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F3/0613G06F3/0659G06F3/0679G06F3/1438G06N3/0445G06N3/0454G06N3/063G06N3/08G06N3/084G06T1/60G09G5/363G09G5/001G09G2352/00G09G2360/06G09G2360/08G09G2360/121G09G2360/123G09G2370/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,417,734
App. No.
15/698,217
Granted
Sep 17, 2019
Kind
B2
Abstract

An apparatus to facilitate compute optimization is disclosed. The apparatus includes a memory device including a first integrated circuit (IC) including a plurality of memory channels and a second IC including a plurality of processing units, each coupled to a memory channel in the plurality of memory channels.

Claims (49)

1. An apparatus, comprising:

a graphics processor; and

a memory device coupled with the graphics processor, the memory device including:

a first integrated circuit (IC) including:

a plurality of memory channels,

a plurality of high-speed interfaces, each of the plurality of memory channels being coupled to a respective interface of the plurality of high-speed interfaces, and

a memory;

a second IC including a plurality of processing units, each processing unit of the plurality of processing units being coupled to a respective memory channel of the plurality of memory channels via the respective interface for the memory channel; and

a controller, the controller to facilitate access by the plurality of processing units to the plurality of memory channels to perform compute operations in response to an instruction from the graphics processor, the compute operations to be performed without transfer of data from the memory to the graphics processor.

2. The apparatus of claim 1 , wherein the compute operations include data intensive functions.

3. The apparatus of claim 2 , wherein the compute operations include operations for a deep neural network.

4. The apparatus of claim 1 , wherein the memory device includes:

a first memory channel of the plurality of memory channels;

a first interface of the plurality of high-speed interfaces, the first interface being coupled to the first memory channel; and

a first processing unit of the plurality of processing units coupled to the first interface.

5. The apparatus of claim 4 , wherein the controller is to facilitate access of the first memory channel by the first processing unit via the first interface to perform the compute operations.

6. The apparatus of claim 5 , wherein the controller is further to facilitate storing of results of the compute operations to the memory via the first memory channel by the first processing unit.

7. The apparatus of claim 6 , wherein the controller is further to facilitate access by the graphics processor to the results of the compute operations stored in the memory via the first memory channel.

8. The apparatus of claim 1 , wherein the memory comprises a high bandwidth memory (HBM), each of the plurality of memory channels being coupled to the HBM.

9. An integrated circuit (IC) package, comprising:

a first integrated circuit (IC) coupled with a graphics processor, the first IC including:

a plurality of memory channels,

a plurality of high-speed interfaces, each of the plurality of high-speed interfaces being coupled to a respective memory channel of the plurality of memory channels, and

a memory;

a second IC including a plurality of processing units, each processing unit of the plurality of processing units being coupled to a memory channel of the plurality of memory channels via the respective interface for the memory channel, and

a controller, the controller to facilitate access by the plurality of processing units to the plurality of memory channels to perform compute operations in response to an instruction from the graphics processor, the compute operations to be performed without transfer of data from the memory to the graphics processor.

10. The IC package of claim 9 , wherein the compute operations include data intensive functions.

11. The IC package of claim 10 , wherein the compute operations include operations for a deep neural network.

12. The IC package of claim 9 , wherein the controller is further to facilitate storing of results of the compute operations to the memory via each of the plurality of memory channels by the respective processing units of the plurality of processing units.

13. The IC package of claim 12 , wherein the controller is further to facilitate access by the graphics processor to the results of the compute operations stored in the memory via the respective memory channels of the plurality of memory channels.

14. The IC package of claim 9 , wherein the memory comprises a high bandwidth memory (HBM), each of the plurality of memory channels being coupled to the HBM.

15. A system, comprising:

a graphics processing unit (GPU); and

a memory device coupled with the GPU, the memory device including:

a memory integrated circuit (IC) comprising:

a plurality of memory channels,

a plurality of high speed interfaces, each of the plurality of memory channels being coupled to a respective interface of the plurality of high speed interfaces, and

a high bandwidth memory (HBM), each of the plurality of memory channels being coupled to the HBM;

a processor IC including a plurality of processing units, each processing unit of the plurality of processing units being coupled to a respective memory channel of the plurality of memory channels via the respective interface for the memory channel; and

a controller, the controller to facilitate access by the plurality of processing units to the plurality of memory channels to perform compute operations in response to an instruction from the GPU, the compute operations to be performed without transfer of data from the memory to the GPU.

16. The system of claim 15 , wherein the compute operations include data intensive functions.

17. The system of claim 16 , wherein the compute operations include operations for a deep neural network.

18. The system of claim 15 , further comprising:

a first memory channel of the plurality of memory channels;

a first interface of the plurality of high speed interfaces, the first interface being coupled to the first memory channel; and

a first processing unit of the plurality of processing units coupled to the first interface.

19. The system of claim 18 , wherein the memory device further comprises a controller is to facilitate access of the first memory channel by the first processing unit via the first interface to perform the compute operations.

20. The system of claim 19 , wherein the controller is further to facilitate storing of results of the compute operations to the HBM via the first memory channel by the first processing unit.

21. The system of claim 20 , wherein the controller is further to facilitate access by the GPU to the compute operations results stored in the HBM via the first memory channel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2019
From: SURTI, PRASOONKUMAR; SRINIVASA, NARAYAN; CHEN, FENG; RAY, JOYDEEP; ASHBAUGH, BEN J.; GALOPPO VON BORRIES, NICOLAS C.; NURVITADHI, ERIKO; VEMBU, BALAJI; LIN, TSUNG-HAN; SINHI, KAMAL; BARIK, RAJKISHORE; BAGHSORKHI, SARA S.; GOTTSCHLICH, JUSTIN E.; KOKER, ALTUG; SATISH, NADATHUR RAJAGOPALAN; AKHBARI, FARSHAD; KIM, DUKHWAN; FU, WENYIN; SCHLUESSLER, TRAVIS T.; MASTRONARDE, JOSH B.; HURD, LINDA L.; FEIT, JOHN H.; BOLES, JEFFERY S.; LAKE, ADAM T.; APPU, ABHISHEK R.; VAIDYANATHAN, KARTHIK; BURKE, DEVAN; MAIYURAN, SUBRAMANIAM
To: INTEL CORPORATION
Reel/Frame 049081/0470 →
Continuity (2)
Continuation In Part 15494886 · Apr 24, 2017
Related Publication 20180308206A1 · Oct 25, 2018