IP Library › Granted Patent US 12,524,702
Granted Patent B2
US 12,524,702 · App. 17/149,602 · Granted Jan 13, 2026

Computing dot products at hardware accelerator

Inventors: Derek Edward Davout Gladding (Poughquag, NY); Nitin Naresh Garegrat (San Jose, CA); Viraj Sunil Khadye (Sunnyvale, CA); Yuxuan Zhang (San Jose, CA)
Assignee: Microsoft Technology Licensing, LLC
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,702
App. No.
17/149,602
Granted
Jan 13, 2026
Kind
B2
Abstract

A computing device, including a hardware accelerator configured to train a machine learning model by computing a first product matrix including a plurality of first dot products. Computing the first product matrix may include receiving a first matrix including a plurality of first vectors and a second matrix including a plurality of second vectors. Each first vector may include a first shared exponent and a plurality of first vector elements. Each second vector may include a second shared exponent and a plurality of second vector elements. For each first vector, computing the first product matrix may further include computing the first dot product of the first vector and a second vector. The first dot product may include a first dot product exponent, a first dot product sign, and a first dot product mantissa. Training the first machine learning model may further include storing the first product matrix in memory.

Claims (77)

1 . A computing device comprising:

a hardware accelerator configured to train a machine learning model at least in part by:

computing a first product matrix including a plurality of first dot products, wherein computing the first product matrix includes:

receiving a first matrix including a plurality of first vectors and a second matrix including a plurality of second vectors, wherein:

each first vector of the plurality of first vectors includes a first shared exponent and a plurality of first vector elements; and

each second vector of the plurality of second vectors includes a second shared exponent and a plurality of second vector elements; and

for each first vector of the plurality of first vectors, computing the first dot product of the first vector and a second vector of the plurality of second vectors, wherein the first dot product includes a first dot product exponent, a first dot product sign, and a first dot product mantissa; and

storing the first product matrix in memory, wherein:

the hardware accelerator includes a plurality of pipeline stages that each include a corresponding multiplier block;

the hardware accelerator is configured to compute a corresponding plurality of product matrices, including the first product matrix, at the multiplier blocks of the plurality of pipeline stages; and

when computing the product matrices, two or more pipeline stages of the plurality of pipeline stages are configured to perform respective matrix multiplication operations on input matrices that have a shared-exponent data type and on input matrices that have an unshared-exponent data type.

2 . The computing device of claim 1 , wherein:

each first vector element of the plurality of first vector elements includes a respective first element sign and a respective first element mantissa; and

each second vector element of the plurality of second vector elements includes a respective second element sign and a respective second element mantissa.

3 . The computing device of claim 1 , wherein the hardware accelerator is reconfigurable to compute a second product matrix including a plurality of second dot products at least in part by:

receiving a third matrix including a plurality of third vectors and a fourth matrix including a plurality of fourth vectors, wherein:

each third vector of the plurality of third vectors includes a plurality of third vector elements that each include a respective third element exponent, a respective third element sign, and a respective third element mantissa; and

each fourth vector of the plurality of fourth vectors includes a plurality of fourth vector elements that each include a respective fourth element exponent, a respective fourth element sign, and a respective fourth element mantissa;

for each third vector of the plurality of third vectors, computing the second dot product of the third vector and a fourth vector of the plurality of fourth vectors, wherein the second dot product includes a second dot product exponent, a second dot product sign, and a second dot product mantissa; and

storing the second product matrix in the memory.

4 . The computing device of claim 3 , wherein:

each first vector element of the plurality of first vector elements and each second vector element of the plurality of second vector elements includes a respective mantissa having a first mantissa length; and

each third vector element of the plurality of third vector elements and each fourth vector element of the plurality of fourth vector elements includes a respective mantissa having a second mantissa length, wherein the first mantissa length is different from the second mantissa length.

5 . The computing device of claim 4 , wherein:

the plurality of first dot products and the plurality of second dot products are computed at two or more of the multiplier blocks included in the hardware accelerator;

the second mantissa length is an integer multiple of the first mantissa length; and

the hardware accelerator is further configured to reconfigure the plurality of multiplier blocks to receive the plurality of third vectors and the plurality of fourth vectors at least in part by combining the plurality of multiplier blocks into a multiplier super-block at which the plurality of second dot products are computed.

6 . The computing device of claim 4 , wherein:

the plurality of first dot products and the plurality of second dot products are computed at a multiplier block of the plurality of multiplier blocks included in the hardware accelerator;

the first mantissa length is an integer multiple of the second mantissa length; and

the hardware accelerator is further configured to reconfigure the multiplier block to receive the plurality of third vectors and the plurality of fourth vectors at least in part by dividing the multiplier block into a plurality of multiplier sub-blocks at which the plurality of second dot products are computed.

7 . The computing device of claim 1 , wherein the inputs received at the two or more pipeline stages include respective input type metadata indicating the respective input types of the inputs.

8 . The computing device of claim 1 , wherein computing the first product matrix further includes performing an exponent normalization operation on the first dot product.

9 . The computing device of claim 8 , wherein computing the first product matrix further includes:

adding the first dot product to an additional dot product to obtain a dot product sum; and

performing the exponent normalization operation on the dot product sum.

10 . A computing device comprising:

a hardware accelerator configured to train a machine learning model at least in part by:

computing a first product matrix, wherein computing the first product matrix includes:

configuring a multiplier block to receive inputs that have a shared-exponent data type;

at the multiplier block, receiving a first vector and a second vector that each have the shared-exponent data type;

computing a first dot product of the first vector and the second vector;

computing a second product matrix, wherein computing the second product matrix includes:

reconfiguring the multiplier block to receive inputs that have an unshared-exponent data type;

receiving a third vector and a fourth vector that each have the unshared-exponent data type; and

computing a second dot product of the third vector and the fourth vector; and

storing the first product matrix and the second product matrix in memory, wherein:

the hardware accelerator includes a plurality of pipeline stages;

the multiplier block is included in a pipeline stage of the plurality of pipeline stages; and

two or more pipeline stages of the plurality of pipeline stages are configured to perform respective matrix multiplication operations on input matrices that have the shared-exponent data type and on input matrices that have the unshared-exponent data type.

11 . The computing device of claim 10 , wherein:

each first vector element of a plurality of first vector elements included in the first vector and each second vector element of a plurality of second vector elements included in the second vector includes a respective mantissa having a first mantissa length; and

each third vector element of a plurality of third vector elements included in the third vector and each fourth vector element of a plurality of fourth vector elements included in the fourth vector includes a respective mantissa having a second mantissa length, wherein the first mantissa length is different from the second mantissa length.

12 . The computing device of claim 11 , wherein:

a plurality of first dot products and a plurality of second dot products are computed at a plurality of multiplier blocks included in the hardware accelerator; and

the hardware accelerator is configured to reconfigure the plurality of multiplier blocks to receive inputs that have the unshared-exponent data type at least in part by combining a plurality of multiplier blocks including the multiplier block into a multiplier super-block at which the second dot product is computed.

13 . The computing device of claim 11 , wherein, when the multiplier block is reconfigured to receive inputs that have the unshared-exponent data type, the hardware accelerator is further configured to reconfigure the multiplier block at least in part by dividing the multiplier block into a plurality of multiplier sub-blocks at which the second dot product is computed.

14 . The computing device of claim 10 , wherein computing the first product matrix and the second product matrix further includes performing an exponent normalization operation on the first dot product and the second dot product.

15 . The computing device of claim 14 , wherein the hardware accelerator is further configured to perform the exponent normalization operation on a plurality of intermediate products of third vector elements of the third vector and fourth vector elements of the fourth vector when computing the second dot product.

16 . A method for use with a computing device, the method comprising:

at a hardware accelerator, training a machine learning model at least in part by:

computing a first product matrix including a plurality of first dot products, wherein computing the first product matrix includes:

receiving a first matrix including a plurality of first vectors and a second matrix including a plurality of second vectors, wherein:

each first vector of the plurality of first vectors includes a first shared exponent and a plurality of first vector elements; and

each second vector of the plurality of second vectors includes a second shared exponent and a plurality of second vector elements; and

for each first vector of the plurality of first vectors, computing the first dot product of the first vector and a second vector of the plurality of second vectors, wherein the first dot product includes a first dot product exponent, a first dot product sign, and a first dot product mantissa; and

storing the first product matrix in memory, wherein:

the hardware accelerator includes a plurality of pipeline stages that each include a corresponding multiplier block;

the method further comprises computing a corresponding plurality of product matrices, including the first product matrix, at the multiplier blocks of the plurality of pipeline stages; and

when computing the product matrices, at two or more pipeline stages of the plurality of pipeline stages, the method further comprises performing respective matrix multiplication operations on input matrices that have a shared-exponent data type and on input matrices that have an unshared-exponent data type.

17 . The method of claim 16 , further comprising reconfiguring the hardware accelerator to compute a second product matrix including a plurality of second dot products at least in part by:

receiving a third matrix including a plurality of third vectors and a fourth matrix including a plurality of fourth vectors, wherein:

each third vector of the plurality of third vectors includes a plurality of third vector elements that each include a respective third element exponent, a respective third element sign, and a respective third element mantissa; and

each fourth vector of the plurality of fourth vectors includes a plurality of fourth vector elements that each include a respective fourth element exponent, a respective fourth element sign, and a respective fourth element mantissa;

for each third vector of the plurality of third vectors, computing the second dot product of the third vector and a fourth vector of the plurality of fourth vectors, wherein the second dot product includes a second dot product exponent, a second dot product sign, and a second dot product mantissa; and

storing the second product matrix in the memory.

18 . The method of claim 16 , wherein computing the first product matrix further includes performing an exponent normalization operation on the first dot product.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2021
From: GLADDING, DEREK EDWARD DAVOUT; GAREGRAT, NITIN NARESH; KHADYE, VIRAJ SUNIL; ZHANG, YUXUAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 054927/0132 →
Continuity (1)
Related Publication 20220222575A1 · Jul 14, 2022
References Cited (20)
US 8275822B2 · Olofsson · 2012 [cited by examiner]
US 9104474B2 · Kaul et al. · 2015 [cited by applicant]
US 10042607B2 · Langhammer · 2018 [cited by applicant]
US 10642614B2 · Kaul · 2020 [cited by examiner]
US 20180322607A1 · Mellempudi et al. · 2018 [cited by applicant]
US 20190339937A1 · Lo et al. · 2019 [cited by applicant]
US 20200193273A1 · Chung et al. · 2020 [cited by applicant]
US 20200272881A1 · Lo · 2020 [cited by applicant]
US 20200302330A1 · Chung et al. · 2020 [cited by applicant]
CN 111767516A · 2020 [cited by applicant]
EP 3713093A1 · 2020 [cited by applicant]
Zhang et al., New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference, 2019 (Year : 2019). [cited by examiner]
Luo et al., Accelerating Pipelined Integer and Floating-Point Accumulations in Configurable Hardware with Delayed Addition Techniques, 2000 (Year: 2000). [cited by examiner]
Albericio, et al., “Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing”, In Proceedings of ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Jun. 18, 2016, pp. 1-13. [cited by applicant]
Belyaev, et al., “A High-perfomance Multi-format SIMD Multiplier for Digital Signal Processors”, In Proceedings of EEE Conference of Russian Young Researchers in Electrical and Electronic Engineering (EIConRus), Jan. 27… [cited by applicant]
Condia, et al., “FlexGripPlus: An improved GPGPU model to support reliability analysis”, In Journal of Microelectronics Reliability, Jun. 1, 2020, pp. 1-14. [cited by applicant]
Kuang, et al., “A Multi-functional Multi-precision 4D Dot Product Unit with SIMD Architecture”, In Arabian Journal for Science and Engineering, vol. 41, Mar. 30, 2016. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US21/061302”, Mailed Date: Mar. 7, 2022, 12 Pages. [cited by applicant]
Zhang, et al., “New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference”, In Proceedings of IEEE Transactions on Computers, vol. 69, Issue 1, J, Jan. 1, 2020, pp. 26-38. [cited by applicant]
Office Action Received for Taiwan Application No. 110146499, mailed on Oct. 28, 2025, 11 Pages (English Translation Not Provided). [cited by applicant]