IP Library › Granted Patent US 10,353,706
Granted Patent B2
US 10,353,706 · App. 15/819,152 · Granted Jul 16, 2019

Instructions and logic to perform floating-point and integer operations for machine learning

Inventors: Himanshu Kaul (Portland, OR); Mark A. Anders (Hillsboro, OR); Sanu K. Mathew (Hillsboro, OR); Anbang Yao (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Tatiana Shpeisman (Menlo Park, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Kamal Sinha (Rancho Cordova, CA); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Vasanth Ranganathan (El Dorado Hills, CA); Sanjeev Jahagirdar (Folsom, CA)
Assignee: Intel Corporation
G06F9/3001G06F7/483G06F7/5443G06F9/30014G06F9/30036G06F9/3851G06N3/0445G06N3/0454G06N3/063G06N3/08G09G5/393G06F2207/3824G06N20/00G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,353,706
App. No.
15/819,152
Granted
Jul 16, 2019
Kind
B2
Abstract

One embodiment provides for a graphics processing unit to accelerate machine-learning operations, the graphics processing unit comprising a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction; and a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute a 32-bit intermediate product of 16-bit operands and to compute a 32-bit sum based on the 32-bit intermediate product.

Claims (32)

1. A graphics processing unit to accelerate machine-learning operations, the graphics processing unit comprising:

a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction across multiple threads of the multiprocessor, wherein at least a portion of a register file is dedicated to each of the multiple threads; and

a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute a 32-bit intermediate product of 16-bit operands and to compute a 32-bit sum based on the 32-bit intermediate product, wherein the first compute unit is to compute the 32-bit intermediate product from two or more 16-bit operands of the at least one single instruction, perform a 16-bit floating-point multiply on the 16-bit operands to generate a 16-bit product, and convert the 16-bit product to the 32-bit intermediate product.

2. The graphics processing unit as in claim 1 , the multiprocessor to execute parallel threads of a thread group, each thread of the thread group having independent thread state.

3. The graphics processing unit as in claim 2 , the multiprocessor including a scheduler to schedule the parallel threads of the thread group to multiple compute units within the multiprocessor.

4. The graphics processing unit as in claim 3 , the multiple compute units within the multiprocessor including a second compute unit to perform an integer operation, the scheduler to schedule a floating-point operation to the first compute unit and an integer operation to the second compute unit, wherein the multiprocessor is to concurrently execute a floating-point operation on the first compute unit and an integer operation on the second compute unit.

5. The graphics processing unit as in claim 4 , the multiprocessor to concurrently execute, on the first compute unit, a first floating-point operation at a first precision and a second floating-point operation at a second precision.

6. The graphics processing unit as in claim 1 , the first compute unit additionally including one or more shifters to normalize or align an intermediate result.

7. The graphics processing unit as in claim 6 , the first compute unit to compute a 16-bit sum based on the 32-bit intermediate product.

8. A data processing system comprising:

a graphics processing unit to accelerate machine-learning operations, the graphics processing unit including a multiprocessor having a single instruction, multiple thread (SIMT) architecture, the multiprocessor to execute at least one single instruction across multiple threads of the multiprocessor, wherein at least a portion of a register file is dedicated to each of the multiple threads;

a first compute unit included within the multiprocessor, the at least one single instruction to cause the first compute unit to perform a two-dimensional matrix multiply and accumulate operation, wherein to perform the two-dimensional matrix multiply and accumulate operation includes to compute a 32-bit intermediate product of 16-bit operands and to compute a 32-bit sum based on the 32-bit intermediate product, wherein the first compute unit is to compute the 32-bit intermediate product from two or more 16-bit operands of the at least one single instruction, perform a 16-bit floating-point multiply on the 16-bit operands to generate a 16-bit product, and convert the 16-bit product to the 32-bit intermediate product; and

a memory communicatively coupled with the graphics processing unit.

9. The data processing system as in claim 8 , the multiprocessor to execute parallel threads of a thread group, each thread of the thread group having independent thread state.

10. The data processing system as in claim 9 , the multiprocessor including a scheduler to schedule the parallel threads to multiple compute units within the multiprocessor.

11. The graphics processing unit as in claim 10 , the multiple compute units within the multiprocessor including a second compute unit to perform an integer operation, the scheduler to schedule a floating-point operation to the first compute unit and an integer operation to the second compute unit, wherein the multiprocessor is to concurrently execute a floating-point operation on the first compute unit and an integer operation on the second compute unit.

12. The graphics processing unit as in claim 11 , the multiprocessor to concurrently execute, on the first compute unit, a first floating-point operation at a first precision and a second floating-point operation at a second precision.

13. The graphics processing unit as in claim 8 , the first compute unit additionally including one or more shifters to normalize or align an intermediate result.

14. The graphics processing unit as in claim 13 , the first compute unit to compute a 16-bit sum based on the 32-bit intermediate product.

15. A method of accelerating a machine-learning operation, the method comprising:

decoding a single instruction on a graphics processing unit (GPU), the GPU having a single instruction, multiple thread (SIMT) architecture;

executing the single instruction via a multiprocessor within the GPU, the single instruction executed across multiple threads of the multiprocessor, wherein at least a portion of a register file is dedicated to each of the multiple threads; and

in response to executing the single instruction via the multiprocessor, performing a two-dimensional matrix multiply and accumulate operation on a first compute unit of the multiprocessor, wherein performing the two-dimensional matrix multiply and accumulate operation includes computing a 32-bit intermediate product of 16-bit operands and computing a 32-bit sum based on the 32-bit intermediate product, wherein computing the 32-bit intermediate product includes performing a 16-bit floating-point multiply on two or more 16-bit operands of the single instruction to generate a 16-bit product and converting the 16-bit product to the 32-bit intermediate product.

16. The method as in claim 15 , additionally comprising executing parallel threads of a thread group, each thread of the thread group having independent thread state.

17. The method as in claim 16 , additionally comprising scheduling the parallel threads of the thread group to multiple compute units within the multiprocessor.

18. The method as in claim 17 , additionally comprising:

scheduling a floating-point operation to the first compute unit and an integer operation to a second compute unit; and

performing the integer operation via a second compute unit within the multiprocessor concurrently with the floating-point operation on the first compute unit.

19. The method as in claim 18 , additionally comprising:

concurrently executing, on the first compute unit, a first floating-point operation at a first precision and a second floating-point operation at a second precision.

20. The method as in claim 15 , additionally comprising:

computing a 16-bit sum based on the 32-bit intermediate product.

Continuity (3)
Continuation 15787129 · Oct 18, 2017
Provisional Application 62491699 · Apr 28, 2017
Related Publication 20180315399A1 · Nov 1, 2018
Cited By (21)
US 12,198,222 US 12,204,487 US 12,210,477 US 12,217,053 US 12,242,414 US 12,293,431 US 12,321,310 US 12,361,600 US 12,386,779 US 12,411,695 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,572,997 US 12,619,376 US 12,670,121 US 12,688,146 US 12,730,759 US 12,737,317 US 12,737,318