IP Library › Granted Patent US 10,748,323
Granted Patent B2
US 10,748,323 · App. 16/208,632 · Granted Aug 18, 2020

GPU based shader constant folding

Inventors: John Gierach (Portland, OR); Srividya Karumuri (Bangalore, IN); Thomas Raoux (Mountain View, CA); Devan Burke (Portland, OR); Wojtek Rajski (Tigard, OR); Jeremy Brennan (Hillsboro, OR)
Assignee: INTEL CORPORATION
G06T15/005G06T1/20G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,748,323
App. No.
16/208,632
Granted
Aug 18, 2020
Kind
B2
Abstract

Embodiments described herein provide a general purpose graphics processing device, comprising a general purpose graphics processing compute block to process a workload including graphics or compute operations, a memory, and a constant folding unit comprising a processing unit to receive a first input shader and metadata for the first input shader, receive a first constant buffer comprising runtime constants for the first input shader, and generate an improved shader from the first input shader and the runtime constants. Other embodiments may be described and claimed.

Claims (63)

1. A general purpose graphics processing device, comprising:

a general purpose graphics processing compute block to process a workload including graphics or compute operations, the general purpose graphics processing compute block comprising a first plurality of processing units comprising one or more registers;

a shared memory communicatively coupled to the first plurality of processing units; and

a constant folding unit comprising an intermediate storage block and a second plurality of processing units, at least one processing unit of the second plurality of processing units to:

receive a first input shader comprising shader code for execution on the general purpose graphics processing compute block and metadata for the first input shader;

load the shader code and metadata into the intermediate storage block;

load the metadata from the intermediate storage block to the one or more registers of the plurality of processing units prior to execution of the first input shader;

receive a first constant buffer comprising runtime constants for the first input shader; and

remove at least a portion of the shader code from the first input shader to generate an improved shader from the first input shader and the runtime constants.

2. The general purpose graphics processing device of claim 1 , the processing unit to:

load a constant fold pass from the memory; and

execute the constant folding pass.

3. The general purpose graphics processing device of claim 2 , the processing unit to:

read a current version of input shader instructions and metadata associated with the first input shader from intermediate storage;

fold one or more constant values into the input shader instructions;

apply one or more pass-specific optimizations to the first input shader; and

write an optimized version of the first input shader to intermediate storage in the memory.

4. The general purpose graphics processing device of claim 3 , the processing unit to:

generate a state programming hardware command associated with the optimized version of the first input shader.

5. The general purpose graphics processing device of claim 1 , wherein the general-purpose graphics processing compute block includes multiple compute clusters, each compute cluster including multiple graphics multiprocessors.

6. The general purpose graphics processing device of claim 1 , wherein the general-purpose graphics processing device is an add-in card connected to the general-purpose processor via a system bus.

7. A heterogeneous data processing system comprising:

a general purpose processor; and

a general purpose graphics processing device comprising:

a general purpose graphics processing compute block to process a workload including graphics or compute operations, the general purpose graphics processing compute block comprising a first plurality of processing units comprising one or more registers;

a shared memory communicatively coupled to the first plurality of processing units; and

a constant folding unit comprising an intermediate storage block and a second plurality of processing units, at least one processing unit of the second plurality of processing units to:

receive a first input shader comprising shader code for execution on the general purpose graphics processing compute block and metadata for the first input shader;

load the metadata from the intermediate storage block to the one or more registers of the plurality of processing units prior to execution of the first input shader;

receive a first constant buffer comprising runtime constants for the first input shader; and

remove at least a portion of the shader code from the first input shader to generate an improved shader from the first input shader and the runtime constants.

8. The heterogeneous data processing system of claim 7 , the processing unit to:

load a constant fold pass from the memory; and

execute the constant folding pass.

9. The heterogeneous data processing system of claim 7 , the processing unit to:

read a current version of input shader instructions and metadata associated with the first input shader from intermediate storage;

fold one or more constant values into the input shader instructions;

apply one or more pass-specific optimizations to the first input shader; and

write an optimized version of the first input shader to intermediate storage in the memory.

10. The heterogeneous data processing system claim 9 , the processing unit to:

generate a state programming hardware command associated with the optimized version of the first input shader.

11. The heterogeneous data processing system of claim 7 , wherein the general-purpose graphics processing compute block includes multiple compute clusters, each compute cluster including multiple graphics multiprocessors.

12. The heterogeneous data processing system of claim 7 , wherein the general-purpose graphics processing device is an add-in card connected to the separate general-purpose processor via a system bus.

13. A non-transitory computer readable medium comprising instructions which, when executed by a general purpose graphics processing device, cause the general purpose graphics processing device to:

process a workload including graphics or compute operations on a general purpose graphics processing compute block of the general purpose graphics processing device, the general purpose graphics processing compute block comprising a first plurality of processing units comprising one or more registers; and

in a constant folding unit comprising a second plurality of processing units, at least one processing unit of the second plurality of processing units:

receive a first input shader comprising shader code for execution on the general purpose graphics processing compute block and metadata for the first input shader;

load the shader code and metadata into the intermediate storage block;

load the metadata from the intermediate storage block to the one or more registers of the plurality of processing units prior to execution of the first input shader;

receive a first constant buffer comprising runtime constants for the first input shader; and

remove at least a portion of the shader code from the first input shader to generate an improved shader from the first input shader and the runtime constants.

14. The non-transitory computer readable medium of claim 13 , the constant folding unit further to:

load a constant fold pass from the memory; and

execute the constant folding pass.

15. The non-transitory computer readable medium of claim 13 , the constant folding unit further to:

read a current version of input shader instructions and metadata associated with the first input shader from intermediate storage;

fold one or more constant values into the input shader instructions;

apply one or more pass-specific optimizations to the first input shader; and

write an optimized version of the first input shader to intermediate storage in the memory.

16. The non-transitory computer readable medium of claim 15 , the constant folding unit further to:

generate a state programming hardware command associated with the optimized version of the first input shader.

17. The non-transitory computer readable medium of claim 13 , wherein the general-purpose graphics processing compute block includes multiple compute clusters, each compute cluster including multiple graphics multiprocessors.

18. The non-transitory computer readable medium of claim 13 , wherein the general-purpose graphics processing device is an add-in card connected to the separate general-purpose processor via a system bus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2019
From: GIERACH, JOHN; KARUMURI, SRIVIDYA; RAOUX, THOMAS; BURKE, DEVAN; RAJSKI, WOJTEK; BRENNAN, JEREMY
To: INTEL CORPORATION
Reel/Frame 047958/0379 →
Continuity (1)
Related Publication 20200175741A1 · Jun 4, 2020