IP Library Granted Patent US 10,489,877
Granted Patent B2
US 10,489,877 · App. 15/494,905 · Granted Nov 26, 2019

Compute optimization mechanism

Inventors: Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Linda L. Hurd (Cool, CA); Dukhwan Kim (San Jose, CA); Mike B. Macpherson (Portland, OR); John C. Weast (Portland, OR); Feng Chen (Shanghai, CN); Farshad Akhbari (Chandler, AZ); Narayan Srinivasa (Portland, OR); Nadathur Rajagopalan Satish (Santa Clara, CA); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Anbang Yao (Beijing, CN); Tatiana Shpeisman (Menlo Park, CA)
Assignee: Intel Corporation
G06T1/20G06F3/14G06F9/3001G06F9/3017G06F9/3887G06F9/3895G06N3/0445G06N3/0454G06N3/063G06N3/084G06T15/005G09G5/363G06F9/3851G06T15/04G09G2360/06G09G2360/08G09G2360/121
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,489,877
App. No.
15/494,905
Granted
Nov 26, 2019
Kind
B2
Abstract

An apparatus to facilitate compute optimization is disclosed. The apparatus includes sorting logic to sort processing threads into thread groups based on bit depth of floating point thread operations.

Claims (32)

1. A graphics processor comprising:

a graphics memory device;

a memory controller coupled to the graphics memory device; and

a compute unit having a multi-issue, multi-threaded architecture, the compute unit coupled to the memory controller and the graphics memory device, wherein the compute unit is to perform compute operations at multiple precisions, wherein the compute unit is to execute a set of compute operations associated with a plurality of threads, the set of compute operations including operations at multiple precisions, wherein to execute the set of compute operations, the compute unit is to sort the set of compute operations into a first bin or a second bin, and wherein the first bin is associated with a first precision and the second bin associated with a second precision.

2. The graphics processor as in claim 1 , further comprising a cache memory coupled with the memory controller, the graphics memory device, and the compute unit.

3. The graphics processor as in claim 2 , wherein the cache memory is to perform a load operation to load operands of the set of compute operations from the graphics memory device.

4. The graphics processor as in claim 1 , wherein the compute unit includes a first floating point unit including hardware to perform 8-bit floating point operations and a second floating point unit including hardware to perform 8-bit floating point operations.

5. The graphics processor as in claim 4 , wherein the set of compute operations includes a first operation including an 8-bit floating point operation, a second operation including an 8-bit floating point operation, and a third operation including a 16-bit floating point operation.

6. The graphics processor as in claim 5 , wherein the first precision is an 8-bit floating point precision and the second precision is a 16-bit floating point precision, the compute unit is to sort the first operation and the second operation to the first bin, and the compute unit is to sort the third operation into the second bin.

7. The graphics processor as in claim 6 , wherein the compute unit is to perform the first operation via the first floating point unit and the second operation via the second floating point unit.

8. The graphics processor as in claim 7 , wherein the compute unit is to perform the third operation via the first floating point unit and the second floating point unit.

9. A method comprising:

receiving a plurality of threads for processing at a compute unit having a multi-issue, multi-threaded architecture, the plurality of threads associated with a set of compute operations to be performed at multiple precisions;

sorting the set of compute operations of the plurality of threads into multiple bins, wherein the multiple bins include a first bin associated with a first precision and a second bin associated with a second precision;

forwarding a first operation and a second operation from the first bin to one or more floating point units of the compute unit for execution at the first precision;

forwarding a third operation from the second bin to one or more floating point units of the compute unit for execution at the second precision; and

executing the first operation and the second operation via the one or more floating point units of the compute unit.

10. The method as in claim 9 , wherein the one or more floating point units include a first floating point unit to perform 8-bit floating point operations and a second floating point unit to perform 8-bit floating point operations.

11. The method as in claim 10 , wherein the first operation includes an 8-bit floating point operation, the second operation includes an 8-bit floating point operation, and the first precision is an 8-bit floating-point precision.

12. The method as in claim 11 , wherein the third operation is a 16-bit floating point precision and the second precision is a 16-bit floating point precision.

13. The method as in claim 12 , additionally comprising performing the first operation via the first floating point unit and the performing the second operation via the second floating point unit.

14. The method as in claim 13 , additionally comprising performing the third operation via the first floating point unit and the second floating point unit.

15. A graphics processing system comprising:

a graphics memory device;

a memory controller coupled to the graphics memory device;

a cache memory coupled with the memory controller and the graphics memory device; and

a compute unit having a multi-issue, multi-threaded architecture, the compute unit coupled to the memory controller, the cache memory, and the graphics memory device, wherein the compute unit is to perform compute operations at multiple precisions, wherein the compute unit is to execute a set of compute operations associated with a plurality of threads, the set of compute operations including operations at multiple precisions, wherein to execute the set of compute operations, the compute unit is to sort the set of compute operations into a first bin or a second bin, and wherein the first bin is associated with a first precision and the second bin associated with a second precision.

16. The graphics processing system as in claim 15 , wherein the cache memory is to perform a load operation to load operands of the set of compute operations from the graphics memory device.

17. The graphics processing system as in claim 15 , wherein the compute unit includes a first floating point unit including hardware to perform 8-bit floating point operations and a second floating point unit including hardware to perform 8-bit floating point operations.

18. The graphics processing system as in claim 17 , wherein the set of compute operations includes a first operation including an 8-bit floating point operation, a second operation including an 8-bit floating point operation, and a third operation including a 16-bit floating point operation, wherein the first precision is an 8-bit floating point precision and the second precision is a 16-bit floating point precision, and wherein the compute unit is to sort the first operation and the second operation to the first bin, and the compute unit is to sort the third operation into the second bin.

19. The graphics processing system as in claim 18 , wherein the compute unit is to perform the first operation via the first floating point unit and the second operation via the second floating point unit.

20. The graphics processing system as in claim 19 , wherein the compute unit is to perform the third operation via the first floating point unit and the second floating point unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2017
From: APPU, ABHISHEK R.; KOKER, ALTUG; HURD, LINDA L.; KIM, DUKHWAN; MACPHERSON, MIKE B.; WEAST, JOHN C.; CHEN, FENG; AKHBARI, FARSHAD; SRINIVASA, NARAYAN; SATISH, NADATHUR RAJAGOPALAN; RAY, JOYDEEP; TANG, PING T.; STRICKLAND, MICHAEL S.; CHEN, XIAOMING; YAO, ANBANG; SHPEISMAN, TATIANA
To: INTEL CORPORATION
Reel/Frame 043991/0543 →
Continuity (1)
Related Publication 20180308201A1 · Oct 25, 2018