IP Library Granted Patent US 10,242,423
Granted Patent B2
US 10,242,423 · App. 15/789,565 · Granted Mar 26, 2019

Compute optimizations for low precision machine learning operations

Inventors: Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Sara S. Baghsorkhi (San Jose, CA); Anbang Yao (Beijing, CN); Kevin Nealis (San Jose, CA); Xiaoming Chen (Shanghai, CN); Altug Koker (El Dorado Hills, CA); Abhishek R. Appu (El Dorado Hills, CA); John C. Weast (Portland, OR); Mike B. Macpherson (Portland, OR); Dukhwan Kim (San Jose, CA); Linda L. Hurd (Cool, CA); Ben J. Ashbaugh (Folsom, CA); Barath Lakshmanan (Chandler, AZ); Liwei Ma (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA)
Assignee: Intel Corporation
G06T1/20G06F7/483G06N99/005G06F3/14G06T1/60G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,242,423
App. No.
15/789,565
Granted
Mar 26, 2019
Kind
B2
Abstract

One embodiment provides an accelerator module comprising a memory stack including multiple memory dies; a graphics processing unit (GPU) coupled with the memory stack via one or more memory controllers, the GPU including a plurality of multiprocessors having a single instruction, multiple thread (SIMT) architecture, the multiprocessors to execute at least one single instruction; the at least one single instruction to cause at least a portion of the GPU to perform a floating-point operation on input having differing precisions; and the floating-point operation is a two-dimensional matrix multiply and accumulate operation.

Claims (26)

1. An accelerator on a multi-chip module, the accelerator comprising:

a memory stack including multiple memory dies; and

a graphics processing unit (GPU) coupled with the memory stack via one or more memory controllers, the GPU including a plurality of multiprocessors having a single instruction, multiple thread (SIMT) architecture, the multiprocessors to execute at least one single instruction, the at least one single instruction to accelerate a linear algebra subprogram associated with a machine learning framework;

the at least one single instruction to cause at least a portion of the GPU to perform a floating-point operation on input having differing precisions, the floating-point operation a two-dimensional matrix multiply and accumulate operation;

wherein at least a portion of the plurality of multiprocessors include a mixed precision core, the mixed precision core to execute a thread of the at least one single instruction, the mixed precision core including a floating-point unit to perform a first operation of the thread at a first precision and a second operation of the thread at a second precision; and

wherein the first operation is a multiply having at least one 16-bit floating-point input and the second operation is an accumulate having a 32-bit floating-point input.

2. The accelerator as in claim 1 , the memory stack including high bandwidth memory.

3. The accelerator as in claim 1 , wherein the memory stack is located on a same physical package as the GPU.

4. The accelerator as in claim 1 , the mixed precision core to perform the first operation at a 16-bit precision and the second operation at a 32-bit precision.

5. The accelerator as in claim 1 , wherein, the first operation has two or more 16-bit floating-point inputs.

6. The accelerator as in claim 1 , the mixed precision core configurable to output a 16-bit floating-point value from the two-dimensional matrix multiply and accumulate operation.

7. A method of accelerating a machine-learning operation, the method comprising:

decoding a single instruction on a graphics processing unit (GPU), the GPU having a single instruction, multiple thread (SIMT) architecture, the GPU coupled with a memory stack via one or more memory controllers; and

executing the single instruction via one or more multiprocessors within the GPU, the single instruction to cause at least a portion of the GPU to perform a two-dimensional matrix multiply and accumulate operation to accelerate a linear algebra subprogram associated with a machine learning framework, wherein executing the single instruction includes executing a thread of the single instruction on a mixed precision core of the one or more multiprocessors, the mixed precision core including a floating-point unit to perform a first operation of the thread at a first precision and a second operation of the thread at a second precision, wherein the first operation is a multiply having at least one 16-bit floating-point input and the second operation is an accumulate having a 32-bit floating-point input.

8. The method as in claim 7 , additionally comprising performing multiple operations on input using the mixed precision core to generate a two-dimensional output matrix and storing the two-dimensional output matrix to the memory stack via the one or more memory controllers.

9. The method as in claim 8 , wherein the memory stack includes high bandwidth memory and is located on a same physical package as the GPU.

10. The method as in claim 7 , wherein the first precision is a 16-bit precision and the second precision is a 32-bit precision.

11. The method as in claim 7 , additionally comprising generating a 16-bit floating-point output of the two-dimensional matrix multiply and accumulate operation.

12. A data processing system comprising:

a memory stack including multiple memory dies, the memory stack including high bandwidth memory; and

a graphics processing unit (GPU) coupled with the memory stack via one or more memory controllers, the GPU including a plurality of multiprocessors having a single instruction, multiple thread (SIMT) architecture, wherein at least a portion of the plurality of multiprocessors include a mixed precision core, the mixed precision core to execute a thread of at least one single instruction, the at least one single instruction to accelerate a linear algebra subprogram associated with a machine learning framework, the mixed precision core including a floating-point unit to perform a first operation of the thread at a first precision and a second operation of the thread at a second precision; and

wherein the at least one single instruction is to cause the mixed precision core to perform a thread of a two-dimensional matrix multiply and accumulate floating-point operation on input having differing precisions, the two-dimensional matrix multiply and accumulate operation including the first operation and the second operation, the first operation a multiply having at least one 16-bit floating-point input and the second operation an accumulate having a 32-bit floating-point input.

13. The data processing system as in claim 12 , the memory stack located on a same physical package as the GPU.

14. The data processing system as in claim 12 , the mixed precision core to perform the first operation at a 16-bit precision and the second operation at a 32-bit precision.

15. The data processing system as in claim 14 , wherein, the first operation has two or more 16-bit floating-point inputs.

16. The data processing system as in claim 12 , the mixed precision core configurable to output a 16-bit floating-point value from the two-dimensional matrix multiply and accumulate operation.

Continuity (2)
Continuation 15581167 · Apr 28, 2017
Related Publication 20180315159A1 · Nov 1, 2018
Cited By (1)
US 12,417,380