IP Library Granted Patent US 11,222,392
Granted Patent B2
US 11,222,392 · App. 16/531,763 · Granted Jan 11, 2022

Compute optimization mechanism for deep neural networks

Inventors: Prasoonkumar Surti (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Joydeep Ray (Folsom, CA); Ben J. Ashbaugh (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Rajkishore Barik (Santa Clara, CA); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Altug Koker (El Dorado Hills, CA); Nadathur Rajagopalan Satish (Santa Clara, CA); Farshad Akhbari (Chandler, AZ); Dukhwan Kim (San Jose, CA); Wenyin Fu (Folsom, CA); Travis T. Schluessler (Hillsboro, OR); Josh B. Mastronarde (Sacramento, CA); Linda L. Hurd (Cool, CA); John H. Feit (Folsom, CA); Jeffery S. Boles (Folsom, CA); Adam T. Lake (Portland, OR); Karthik Vaidyanathan (Berkeley, CA); Devan Burke (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Abhishek R. Appu (El Dorado Hills, CA)
Assignee: Intel Corporation
G06T1/20G06F3/0613G06F3/0659G06F3/0679G06F3/1438G06N3/0445G06N3/0454G06N3/063G06N3/08G06N3/084G06T1/60G09G5/363G09G5/001G09G2352/00G09G2360/06G09G2360/08G09G2360/121G09G2360/123G09G2370/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,222,392
App. No.
16/531,763
Filed
Aug 5, 2019
Granted
Jan 11, 2022
Kind
B2
Art Unit
2612
USPC
345/505
Abstract

An apparatus to facilitate compute optimization is disclosed. The apparatus includes a memory device including a first integrated circuit (IC) including a plurality of memory channels and a second IC including a plurality of processing units, each coupled to a memory channel in the plurality of memory channels.

Claims (39)

1. An apparatus, comprising:

one or more processors including a graphics processor, the graphics processor including:

one or more processing units to provide a first set of shader operations associated with a shader stage of a graphics pipeline,

a scheduler to schedule shader threads for processing, and

a field-programmable gate array (FPGA) dynamically configured to provide a second set of shader operations associated with the shader stage of the graphics pipeline; and

wherein, upon the FPGA being dynamically configured to provide the second set of shader operations, one or more shader threads are scheduled to be processed at either the one or more processing units or the FPGA based on one or more processing characteristics targeted for each shader thread, the FPGA being asynchronously triggered to perform operations associated with the shader stage in the graphics pipeline via access to a memory location associated with the FPGA.

2. The apparatus of claim 1 , wherein the first set of shader operations have a first set of processing characteristics and the second set of shader operations have a second set of processing characteristics.

3. The apparatus of claim 2 , wherein the first set of processing characteristics include a first target speed of operation and a first target power consumption and a first shader thread is to be scheduled to be processed by the one or more processing units according to the first set of processing characteristics, and wherein the second set of processing characteristics include a second speed of operation and a second power consumption and a second shader thread is scheduled to be processed at the second set of shader operations according to the second set of processing characteristics.

4. The apparatus of claim 1 , wherein the first set of shader operations and the second set of shader operations are different versions of a same set of shader operations.

5. The apparatus of claim 1 , wherein the graphics processor further includes a buffer to store data for shader operations; and

wherein the one or more processors are to dispatch a request to the FPGA to perform a shader operation in the second set of shader operations via the buffer, the FPGA is to perform the shader operation, and the FPGA is to return a signal to the buffer upon completion of the shader operation.

6. The apparatus as in claim 1 , wherein the one or more processing units include one or more streaming multiprocessors having a single instruction multiple thread architecture.

7. A method comprising:

receiving a plurality of shader threads at a graphics processor, the graphics processor including one or more processing units to process data and a field-programmable gate array (FPGA), wherein the graphics processor provides a graphics pipeline and the one or more processing units are configured to provide a first set of shader operations having a first set of processing characteristics;

dynamically configuring the FPGA as a shader stage in the graphics pipeline, wherein dynamically configuring the FPGA as the shader stage in the graphics pipeline includes asynchronously triggering the FPGA via access to a memory location associated with the FPGA, the FPGA to provide a second set of shader operations, the second set of shader operations having a second set of processing characteristics; and

scheduling one or more shader threads for processing at the one or more processing units or the FPGA based on processing characteristics targeted for each of the one or more shader threads.

8. The method as in claim 7 , wherein the first set of processing characteristics include a first speed of operation and a first power consumption and a first shader thread is scheduled to be processed by the one or more processing units according to the first set of processing characteristics, and wherein the second set of processing characteristics include a second speed of operation and a second power consumption and a second shader thread is scheduled to be processed at the second set of shader operations according to the second set of processing characteristics.

9. The method as in claim 7 , wherein the first set of shader operations and the second set of shader operations are different versions of a same set of shader operations.

10. The method as in claim 7 , wherein the graphics processor includes a buffer to store data for shader operations and the method further comprises dispatching a request to the FPGA to perform a shader operation in the second set of shader operations via the buffer, the FPGA to perform the shader operation and return a signal to the buffer upon completion of the shader operation.

11. The method as in claim 7 , wherein the one or more processing units include one or more streaming multiprocessors having a single instruction multiple thread architecture.

12. A graphics processor, comprising:

one or more processing units to provide a first set of shader operations associated with a shader stage of a graphics pipeline,

a scheduler to schedule shader threads for processing, and

a field-programmable gate array (FPGA) dynamically configured to provide a second set of shader operations associated with the shader stage of the graphics pipeline; and

wherein, upon the FPGA being dynamically configured to provide the second set of shader operations, one or more shader threads are scheduled to be processed at either the one or more processing units or the FPGA based on one or more processing characteristics targeted for each shader thread, the FPGA being asynchronously triggered to perform operations associated with the shader stage in the graphics pipeline via access to a memory location associated with the FPGA.

13. The graphics processor of claim 12 , wherein the first set of shader operations have a first set of processing characteristics, the second set of shader operations have a second set of processing characteristics, and the first set of shader operations and the second set of shader operations are different versions of a same set of shader operations.

14. The graphics processor of claim 13 , wherein the first set of processing characteristics include a first target speed of operation and a first target power consumption a first shader thread is scheduled to be processed by the one or more processing units according to the first set of processing characteristics and the second set of processing characteristics include a second speed of operation and a second power consumption and a second shader thread is scheduled to be processed at the second set of shader operations according to the second set of processing characteristics.

15. The graphics processor of claim 12 , wherein the graphics processor further includes a buffer to store data for shader operations; and

wherein the one or more processing units are to dispatch a request to the FPGA to perform a shader operation in the second set of shader operations via the buffer, the FPGA is to perform the shader operation, and the FPGA is to return a signal to the buffer upon completion of the shader operation.

16. The graphics processor of claim 12 , wherein the one or more processing units include one or more streaming multiprocessors having a single instruction multiple thread architecture.

17. A data processing system comprising:

a memory device;

one or more processors coupled with the memory device, the one or more processors including a graphics processor, the graphics processor including:

a buffer to store data for shader operations;

one or more processing units to provide a first set of shader operations associated with a shader stage of a graphics pipeline,

a scheduler to schedule shader threads for processing, and

a field-programmable gate array (FPGA) dynamically configured to provide a second set of shader operations associated with the shader stage of the graphics pipeline;

wherein, upon the FPGA being dynamically configured to provide the second set of shader operations, one or more shader threads are scheduled to be processed at either the one or more processing units or the FPGA based on one or more processing characteristics targeted for each shader thread, the FPGA being asynchronously triggered to perform operations associated with the shader stage in the graphics pipeline via access to a memory location associated with the FPGA; and

wherein the one or more processors are to dispatch a request to the FPGA to perform a shader operation in the second set of shader operations via the buffer, the FPGA is to perform the shader operation, and the FPGA is to return a signal to the buffer upon completion of the shader operation.

Continuity (3)
Continuation 15698217 · Sep 7, 2017
Continuation In Part 15494886 · Apr 24, 2017
Related Publication 20200034946A1 · Jan 30, 2020
Cited By (1)
US 12,511,163