IP Library Granted Patent US 10,255,656
Granted Patent B2
US 10,255,656 · App. 15/798,574 · Granted Apr 9, 2019

Compute optimization mechanism

Inventors: Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Linda L. Hurd (Cool, CA); Dukhwan Kim (San Jose, CA); Mike B. Macpherson (Portland, OR); John C. Weast (Portland, OR); Feng Chen (Shanghai, CN); Farshad Akhbari (Chandler, AZ); Narayan Srinivasa (Portland, OR); Nadathur Rajagopalan Satish (Santa Clara, CA); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Anbang Yao (Beijing, CN); Tatiana Shpeisman (Menlo Park, CA)
Assignee: INTEL CORPORATION
G06T1/20G06F9/3851G06T15/005G06T15/04G09G5/363
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,255,656
App. No.
15/798,574
Granted
Apr 9, 2019
Kind
B2
Abstract

An apparatus to facilitate compute optimization is disclosed. The apparatus includes sorting logic to sort processing threads into thread groups based on bit depth of floating point thread operations.

Claims (26)

1. A multiprocessor comprising:

a register file to store operands; and

a plurality of processing cores, each core having execution logic to perform mixed precision multi-dimensional matrix fused multiply-accumulate (FMAC) operations, the execution logic to execute one or more instructions to:

multiply a first 16-bit floating point (FP16) operand with a second FP16 operand to obtain a first 32-bit floating point (FP32) intermediate product;

multiply a third FP16 operand with a fourth FP16 operand to obtain a second FP32 intermediate product; and

add the second FP32 intermediate product with the first FP32 intermediate product to generate a FP32 sum result.

2. The multiprocessor of claim 1 , wherein the register file is to store the first FP16 operand, the second FP16 operand, and the first and second FP32 intermediate products.

3. The multiprocessor of claim 1 , further comprising an instruction cache to store the one or more instructions.

4. The multiprocessor of claim 3 , further comprising a dispatch unit to dispatch the one or more instructions for execution at the execution logic.

5. The multiprocessor of claim 1 , further comprising a scheduler to schedule the FMAC operations.

6. A method to facilitate execution of mixed precision multi-dimensional matrix fused multiply-accumulate (FMAC) operations comprising:

receiving a first 16-bit floating point (FP16) operand and a second FP16 operand at one or more processing cores;

multiplying the first FP16 operand with the second FP16 operand to obtain a first 32-bit floating point (FP32) intermediate product;

multiplying a third FP16 operand with a fourth FP16 operand to obtain a second FP32 intermediate product; and

adding the second FP32 intermediate product with the first FP32 intermediate product to generate a FP32 sum result.

7. The method of claim 6 , wherein the first FP16 operand, the second FP16 operand, and the first and second FP32 intermediate products are received from a register file.

8. A graphics processing unit comprising a plurality of multiprocessors, wherein each multiprocessor comprises:

a register file to store operands; and

a plurality of processing cores, each core having execution logic to perform mixed precision multi-dimensional matrix fused multiply-accumulate (FMAC) operations, the execution logic to execute one or more instructions to:

multiply a first 16-bit floating point (FP16) operand with a second FP16 operand to obtain a first 32-bit floating point (FP32) intermediate product;

multiply a third FP16 operand with a fourth FP16 operand to obtain a second FP32 intermediate product; and

add the second intermediate FP32 product with the first FP32 intermediate product to generate a FP32 sum result.

9. The graphics processing unit of claim 8 , wherein the register file is to store the first FP16 operand, the second FP16 operand, and the first and second FP32 intermediate products.

10. The graphics processing unit of claim 8 , further comprising an instruction cache to store the one or more instructions.

11. The graphics processing unit of claim 10 , further comprising a dispatch unit to dispatch the one or more instructions for execution at the execution logic.

12. The graphics processing unit of claim 8 , further comprising a scheduler to schedule the FMAC operations.

Continuity (2)
Continuation 15494905 · Apr 24, 2017
Related Publication 20180308207A1 · Oct 25, 2018
Cited By (1)
US 12,417,380