IP Library › Granted Patent US 11,169,799
Granted Patent B2
US 11,169,799 · App. 16/432,402 · Granted Nov 9, 2021

Instructions and logic to perform floating-point and integer operations for machine learning

Inventors: Himanshu Kaul (Portland, OR); Mark A. Anders (Hillsboro, OR); Sanu K. Mathew (Hillsboro, OR); Anbang Yao (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Tatiana Shpeisman (Menlo Park, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Kamal Sinha (Rancho Cordova, CA); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Vasanth Ranganathan (El Dorado Hills, CA); Sanjeev Jahagirdar (Folsom, CA)
Assignee: Intel Corporation
G06F9/3001G06F7/483G06F7/5443G06F9/30014G06F9/30036G06F9/3851G06N3/0445G06N3/0454G06N3/063G06N3/08G09G5/393G06F2207/3824G06N20/00G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,169,799
App. No.
16/432,402
Filed
Jun 5, 2019
Granted
Nov 9, 2021
Kind
B2
Art Unit
2123
USPC
712/221
Abstract

One embodiment provides for a graphics processing unit to accelerate machine-learning operations, the graphics processing unit comprising a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction; and a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute a 32-bit intermediate product of 16-bit operands and to compute a 32-bit sum based on the 32-bit intermediate product.

Claims (43)

1. A graphics processing unit to accelerate machine-learning operations, the graphics processing unit comprising:

a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction across multiple threads of the multiprocessor; and

a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute an intermediate product of 16-bit operands and to compute a 32-bit sum based on the intermediate product;

wherein to compute a 32-bit sum based on the intermediate product, the first compute unit is to:

perform a floating-point multiply of two or more 16-bit operands at a first intermediate precision to generate the intermediate product, the first intermediate precision less than 32 bits;

compute a sum based on the intermediate product to generate an intermediate sum at a second intermediate precision; and

compute the 32-bit sum via a conversion of the intermediate sum at the second intermediate precision to a 32-bit precision.

2. The graphics processing unit as in claim 1 , the multiprocessor to execute parallel threads of a thread group, each thread of the thread group having independent thread state.

3. The graphics processing unit as in claim 2 , the multiprocessor including a scheduler to schedule the parallel threads of the thread group to multiple compute units within the multiprocessor.

4. The graphics processing unit as in claim 3 , the multiple compute units within the multiprocessor including a second compute unit to perform an integer operation, the scheduler to schedule a floating-point operation to the first compute unit and an integer operation to the second compute unit wherein the multiprocessor is to concurrently execute a floating-point operation on the first compute unit and an integer operation on the second compute unit.

5. The graphics processing unit as in claim 4 , wherein the multiprocessor is to concurrently execute a first floating-point operation at a first precision on the first compute unit and a second floating-point operation at a second precision.

6. The graphics processing unit as in claim 1 , the first compute unit additionally including one or more shifters to normalize or align an intermediate result.

7. The graphics processing unit as in claim 6 , the first compute unit additionally configurable to compute a 16-bit sum based on the intermediate product.

8. A data processing system comprising:

a graphics processing unit to accelerate machine-learning operations, the graphics processing unit including a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction across multiple threads of the multiprocessor; and

a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute an intermediate product of 16-bit operands and to compute a 32-bit sum based on the intermediate product; and

a memory communicatively coupled with the graphics processing unit;

wherein to compute a 32-bit sum based on the intermediate product, the first compute unit is to:

perform a floating-point multiply of two or more 16-bit operands at a first intermediate precision to generate the intermediate product, the first intermediate precision less than 32 bits;

compute a sum based on the intermediate product to generate an intermediate sum at a second intermediate precision; and

compute the 32-bit sum via a conversion of the intermediate sum at the second intermediate precision to a 32-bit precision.

9. The data processing system as in claim 8 , the multiprocessor to execute parallel threads of a thread group, each thread of the thread group having independent thread state.

10. The data processing system as in claim 9 , the multiprocessor including a scheduler to schedule the parallel threads to multiple compute units within the multiprocessor.

11. The graphics processing unit as in claim 10 , the multiple compute units within the multiprocessor including a second compute unit to perform an integer operation, the scheduler to schedule a floating-point operation to the first compute unit and an integer operation to the second compute unit, wherein the multiprocessor is to concurrently execute a floating-point operation on the first compute unit and an integer operation on the second compute unit.

12. The graphics processing unit as in claim 11 , the multiprocessor to concurrently execute, on the first compute unit a first floating-point operation at a first precision and a second floating point operation at a second precision.

13. The graphics processing unit as in claim 8 , the first compute unit additionally including one or more shifters to normalize or align an intermediate result.

14. The graphics processing unit as in claim 13 , the first compute unit to additionally configurable compute a 16-bit sum based on the intermediate product.

15. A method of accelerating a machine-learning operation, the method comprising:

decoding a single instruction on a graphics processing unit (GPU), the GPU having a single instruction, multiple thread (SIMT) architecture;

executing the single instruction via a multiprocessor within the GPU, the single instruction executed across multiple threads of the multiprocessor; and

in response to executing the single instruction via the multiprocessor, performing a two-dimensional matrix multiply and accumulate operation on a first compute unit of the multiprocessor, wherein performing the two-dimensional matrix multiply and accumulate operation includes computing an intermediate product of 16-bit operands and computing a 32-bit sum based on the intermediate product, wherein computing the intermediate product includes:

performing a floating-point multiply of two or more 16-bit operands at a first intermediate precision to generate the intermediate product, the first intermediate precision less than 32-bits;

compute a sum based on the intermediate product to generate an intermediate sum at a second intermediate precision; and

compute the 32-bit sum via a conversion of the intermediate sum at the second intermediate precision to a 32-bit precision.

16. The method as in claim 15 , additionally comprising executing parallel threads of a thread group, each thread of the thread group having independent thread state.

17. The method as in claim 16 , additionally comprising scheduling the parallel threads of the thread group to multiple compute units within the multiprocessor.

18. The method as in claim 17 , additionally comprising:

scheduling a floating-point operation to the first compute unit and an integer operation to a second compute unit; and

performing the integer operation via a second compute unit within the multiprocessor concurrently with the floating-point operation on the first compute unit.

19. The method as in claim 18 , additionally comprising:

concurrently executing, on the first compute unit, a first floating-point operation at a first precision and a second floating-point operation at a second precision.

20. The method as in claim 15 , additionally comprising:

computing a 16-bit sum based on the intermediate product.

Continuity (4)
Continuation 15819152 · Nov 21, 2017
Continuation 15787129 · Oct 18, 2017
Provisional Application 62491699 · Apr 28, 2017
Related Publication 20190369988A1 · Dec 5, 2019
Cited By (20)
US 12,198,222 US 12,204,487 US 12,210,477 US 12,217,053 US 12,242,414 US 12,293,431 US 12,321,310 US 12,361,600 US 12,386,779 US 12,411,695 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,572,997 US 12,670,121 US 12,688,146 US 12,730,759 US 12,737,317 US 12,737,318