IP Library Granted Patent US 10,474,458
Granted Patent B2
US 10,474,458 · App. 15/787,129 · Granted Nov 12, 2019

Instructions and logic to perform floating-point and integer operations for machine learning

Inventors: Himanshu Kaul (Portland, OR); Mark A. Anders (Hillsboro, OR); Sanu K. Mathew (Hillsboro, OR); Anbang Yao (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA); Xiaoming Chen (Shanghai, CN); Tatiana Shpeisman (Menlo Park, CA); Abhishek R. Appu (El Dorado Hills, CA); Altug Koker (El Dorado Hills, CA); Kamal Sinha (Rancho Cordova, CA); Balaji Vembu (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Rajkishore Barik (Santa Clara, CA); Tsung-Han Lin (Campbell, CA); Vasanth Ranganathan (El Dorado Hills, CA); Sanjeev Jahagirdar (Folsom, CA)
Assignee: Intel Corporation
G06F9/3001G06F7/483G06F7/5443G06F9/30014G06F9/30036G06F9/3851G06N3/0445G06N3/0454G06N3/063G06N3/08G09G5/393G06F2207/3824G06N20/00G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,474,458
App. No.
15/787,129
Granted
Nov 12, 2019
Kind
B2
Abstract

One embodiment provides for a machine-learning hardware accelerator comprising a compute unit having an adder and a multiplier that are shared between integer data path and a floating-point datapath, the upper bits of input operands to the multiplier to be gated during floating-point operation.

Claims (35)

1. A method comprising:

fetching and decoding a single instruction to perform a combined multiply and add operation on a set of operands;

issuing the single instruction for execution by a dynamically configurable compute unit;

configuring the dynamically configurable compute unit to perform operations at a precision and data-type of the set of operands; and

executing at least a portion of the single instruction at the dynamically configurable compute unit to generate an output based on the combined multiply and add operation, wherein to generate the output includes quantizing an intermediate value having a first precision to a second precision that is lower than the first precision, the quantizing including stochastically rounding a fractional portion of intermediate data.

2. The method as in claim 1 , wherein the combined multiply and add operation is a fused multiply-add or a fused multiply-accumulate operation.

3. The method as in claim 1 , additionally comprising quantizing the intermediate value via a machine learning accelerator unit.

4. The method as in claim 3 , additionally comprising stochastically rounding the fractional portion of intermediate data based on output of a random number generator.

5. The method as in claim 3 , additionally comprising stochastically rounding the fractional portion of intermediate data based on a probability distribution associated with intermediate data.

6. A hardware accelerator comprising:

a fetch unit to fetch a single instruction to perform a combined multiply and add operation on a set of operands;

a decode unit to decode the single instruction into a decoded instruction; and

a dynamically configurable compute unit to perform operations at a precision and data-type of the set of operands, wherein the dynamically configurable compute unit is to receive the decoded instruction and execute at least a portion of the decoded instruction, wherein to execute at least the portion of the decoded instruction includes to generate an output based on the combined multiply and add operation, to generate the output includes to quantize an intermediate value having a first precision to a second precision that is lower than the first precision, and to quantize the intermediate value includes to stochastically round a fractional portion of intermediate data.

7. The hardware accelerator as in claim 6 , wherein the combined multiply and add operation is a fused multiply-add or a fused multiply-accumulate operation.

8. The hardware accelerator as in claim 6 , wherein the dynamically configurable compute unit is to quantize the intermediate value via a machine learning accelerator unit.

9. The hardware accelerator as in claim 8 , wherein the dynamically configurable compute unit is to stochastically round the fractional portion of intermediate data based on output of a random number generator.

10. The hardware accelerator as in claim 8 , wherein the dynamically configurable compute unit is to stochastically round the fractional portion of intermediate data based on a probability distribution associated with intermediate data.

11. A non-transitory machine-readable medium storing instructions to cause one or more processors of an electronic device to perform operations comprising:

configuring a hardware accelerator to fetch and decode a single instruction to perform a combined multiply and add operation on a set of operands and issue the single instruction for execution by a dynamically configurable compute unit;

configuring the dynamically configurable compute unit to perform operations at a precision and data-type of the set of operands; and

executing at least a portion of the single instruction at the dynamically configurable compute unit to generate an output based on the combined multiply and add operation, wherein to generate the output includes quantizing an intermediate value having a first precision to a second precision that is lower than the first precision, the quantizing including stochastically rounding a fractional portion of intermediate data.

12. The non-transitory machine-readable medium as in claim 11 , wherein the combined multiply and add operation is a fused multiply-add or a fused multiply-accumulate operation.

13. The non-transitory machine-readable medium as in claim 11 , the operations additionally comprising quantizing the intermediate value via the hardware accelerator.

14. The non-transitory machine-readable medium as in claim 13 , the operations additionally comprising stochastically rounding the fractional portion of intermediate data based on output of a random number generator.

15. The non-transitory machine-readable medium as in claim 13 , the operations additionally comprising stochastically rounding the fractional portion of intermediate data based on a probability distribution associated with intermediate data.

16. A data processing system comprising:

a memory device; and

a hardware accelerator coupled with the memory device, wherein the hardware accelerator includes:

a fetch unit to fetch a single instruction to perform a combined multiply and add operation on a set of operands;

a decode unit to decode the single instruction into a decoded instruction; and

a dynamically configurable compute unit to perform operations at a precision and data-type of the set of operands, wherein the dynamically configurable compute unit is to receive the decoded instruction and execute at least a portion of the decoded instruction, wherein to execute at least the portion of the decoded instruction includes to generate an output based on the combined multiply and add operation, to generate the output includes to quantize an intermediate value having a first precision to a second precision that is lower than the first precision, and to quantize the intermediate value includes to stochastically round a fractional portion of intermediate data.

17. The data processing system as in claim 16 , wherein the combined multiply and add operation is a fused multiply-add or a fused multiply-accumulate operation.

18. The data processing system as in claim 16 , wherein the dynamically configurable compute unit is to quantize the intermediate value via a machine learning accelerator unit.

19. The data processing system as in claim 18 , wherein the dynamically configurable compute unit is to stochastically round the fractional portion of intermediate data based on output of a random number generator.

20. The data processing system as in claim 18 , wherein the dynamically configurable compute unit is to stochastically round the fractional portion of intermediate data based on a probability distribution associated with intermediate data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2017
From: KAUL, HIMANSHU; ANDERS, MARK A.; MATHEW, SANU K.; YAO, ANBANG; RAY, JOYDEEP; TANG, PING T.; STRICKLAND, MICHAEL S.; CHEN, XIAOMING; SHPEISMAN, TATIANA; APPU, ABHISHEK R.; KOKER, ALTUG; SINHA, KAMAL; VEMBU, BALAJI; RANGANATHAN, VASANTH; JAHAGIRDAR, SANJEEV; GALOPPO VON BORRIES, NICOLAS C.; NURVITADHI, ERIKO; BARIK, RAJKISHORE; LIN, TSUNG-HAN
To: INTEL CORPORATION
Reel/Frame 043895/0265 →
Continuity (2)
Provisional Application 62491699 · Apr 28, 2017
Related Publication 20180315398A1 · Nov 1, 2018
Cited By (20)
US 12,198,222 US 12,204,487 US 12,210,477 US 12,217,053 US 12,242,414 US 12,293,431 US 12,321,310 US 12,360,740 US 12,361,600 US 12,367,383 US 12,386,779 US 12,411,695 US 12,423,098 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,572,997 US 12,670,121 US 12,688,146