IP Library › Granted Patent US 12,307,250
Granted Patent B2
US 12,307,250 · App. 18/397,664 · Granted May 20, 2025

Systems and methods for performing 16-bit floating-point matrix dot product instructions

Inventors: Alexander F. Heinecke (San Jose, CA); Robert Valentine (Kiryat Tivon, IL); Mark J. Charney (Lexington, MA); Raanan Sade (Portland, OR); Menachem Adelman (Modi'in, IL); Zeev Sperber (Zichron Yackov, IL); Amit Gradstein (Binyamina, IL); Simon Rubanovich (Haifa, IL)
Assignee: Intel Corporation
G06F9/30036G06F9/3001G06F9/30038G06F9/3016G06F9/3802
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,250
App. No.
18/397,664
Granted
May 20, 2025
Kind
B2
Abstract

Disclosed embodiments relate to computing dot products of nibbles in tile operands. In one example, a processor includes decode circuitry to decode a tile dot product instruction having fields for an opcode, a destination identifier to identify a M by N destination matrix, a first source identifier to identify a M by K first source matrix, and a second source identifier to identify a K by N second source matrix, each of the matrices containing doubleword elements, and execution circuitry to execute the decoded instruction to perform a flow K times for each element (m, n) of the specified destination matrix to generate eight products by multiplying each nibble of a doubleword element (M,K) of the specified first source matrix by a corresponding nibble of a doubleword element (K,N) of the specified second source matrix, and to accumulate and saturate the eight products with previous contents of the doubleword element.

Claims (36)

1. A processor comprising:

an instruction cache to store a matrix multiplication instruction;

a plurality of vector registers to store a plurality of single-precision floating point source data elements of a first matrix comprising M rows by N columns, a first plurality of bfloat16 floating-point source data elements of a second matrix comprising M rows and K columns, and a second plurality of bfloat16 floating point source data elements of a third matrix comprising K rows by N columns;

decode circuitry to decode the matrix multiplication instruction, the matrix multiplication instruction including an opcode to indicate a matrix multiplication operation, a first field to indicate a first storage location associated with the plurality of single-precision floating point source data elements, a second field to indicate a second storage location associated with the first plurality of bfloat16 floating-point source data elements, and a third field to indicate a third storage location associated with the second plurality of bfloat16 floating point source data elements, wherein the first, second, and third storage locations are locations in the plurality of vector registers; and

execution circuitry coupled with the decode circuitry, the execution circuitry to perform operations in accordance with the matrix multiplication instruction, the execution circuitry to, for each row m of the M rows of the second matrix and each column n of the N columns of the third matrix:

generate a dot product from K bfloat16 floating-point source data elements corresponding to the row m of the second matrix and K bfloat16 floating-point source data elements corresponding to the column n of the third matrix, and

accumulate the dot product with a single precision floating-point source data element corresponding to a row m of the M rows and a column n of the N columns of the first matrix to generate a single-precision floating-point result data element to be stored in a position of the plurality of vector registers corresponding to the row m and the column n of the first matrix.

2. The processor of claim 1 , wherein the execution circuitry comprises:

a plurality of multipliers to multiply the K bfloat16 floating-point source data elements corresponding to the row m of the second matrix and corresponding K bfloat16 floating-point source data elements corresponding to the column n of the third matrix to generate a plurality of products; and

adder circuitry to add the plurality of products to generate the dot product.

3. The processor of claim 2 , wherein the adder circuitry is additionally to add the dot product and the single precision floating-point source data element.

4. The processor of claim 1 , further comprising: a control register to store an indication of a rounding mode.

5. The processor of claim 4 , wherein the execution circuitry is to perform rounding based on instruction fields of some of the instructions regardless of the indication stored in the control register.

6. The processor of claim 1 , wherein the execution circuitry does not support exceptions due to the matrix multiplication instruction.

7. The processor of claim 1 , wherein the execution circuitry is to perform an operation corresponding to the matrix multiplication instruction to process denormal values of the second matrix and the third matrix as zero.

8. The processor of claim 1 , wherein the execution circuitry is to perform an operation corresponding to the instruction to flush denormal values to zero.

9. The processor of claim 1 , wherein the M rows of the second matrix and the N columns of the third matrix are equal in number.

10. The processor of claim 1 , wherein each dot product is a single precision floating-point dot product.

11. A method comprising:

storing a matrix multiplication instruction to an instruction cache;

storing a plurality of single-precision floating point source data elements of a first matrix comprising M rows by N columns, a first plurality of bfloat16 floating-point source data elements of a second matrix comprising M rows and K columns, and a second plurality of bfloat16 floating point source data elements of a third matrix comprising K rows by N columns in a plurality of vector registers;

decoding the matrix multiplication instruction by decode circuitry, wherein the matrix multiplication instruction includes an opcode to indicate a matrix multiplication operation, a first field to indicate a first storage location associated with the plurality of single-precision floating point source data elements, a second field to indicate a second storage location associated with the first plurality of bfloat 16 floating-point source data elements, and a third field to indicate a third storage location associated with the second plurality of bfloat16 floating point source data elements, wherein the first, second, and third storage locations are locations in the plurality of vector registers; and

executing, by execution circuitry coupled with the decode circuitry, the matrix multiplication instruction, wherein the executing is to, for each row m of the M rows of the second matrix and each column n of the N columns of the third matrix:

generate a dot product from K bfloat 16 floating-point source data elements corresponding to the row m of the second matrix and K bfloat16 floating-point source data elements corresponding to the column n of the third matrix, and

accumulate the dot product with a single precision floating-point source data element corresponding to a row m of the M rows and a column n of the N columns of the first matrix to generate a single-precision floating-point result data element to be stored in a position of the plurality of vector registers corresponding to the row m and the column n of the first matrix.

12. The method of claim 11 , wherein the executing comprises:

multiplying, by a plurality of multipliers of the execution circuitry, the K bfloat16 floating-point source data elements corresponding to the row m of the second matrix and corresponding K bfloat16 floating-point source data elements corresponding to the column n of the third matrix to generate a plurality of products; and

adding, by adder circuitry of the execution circuitry, the plurality of products to generate the dot product.

13. The method of claim 12 , wherein the adder circuitry adds the dot product and the single precision floating-point source data element.

14. The method of claim 11 , further comprising storing an indication of a rounding mode in a control register.

15. The method of claim 14 , wherein the executing by the execution circuitry performs rounding based on instruction fields of some of the instructions regardless of the indication stored in the control register.

16. The method of claim 11 , wherein the execution circuitry does not support exceptions due to the matrix multiplication instruction.

17. The method of claim 11 , wherein the executing by the execution circuitry processes denormal values of the second matrix and the third matrix as zero.

18. The method of claim 11 , the executing by the execution circuitry flushes denormal values to zero.

19. The method of claim 11 , wherein the M rows of the second matrix and the N columns of the third matrix are equal in number.

20. The method of claim 11 , wherein each dot product is a single precision floating-point dot product.

Continuity (4)
Continuation 18190761 · Mar 27, 2023
Continuation 17216566 · Mar 29, 2021
Continuation 16186387 · Nov 9, 2018
Related Publication 20240126545A1 · Apr 18, 2024
References Cited (97)
US 5247632A · Newman · 1993 [cited by applicant]
US 5475822A · Sibigtroth et al. · 1995 [cited by applicant]
US 5892962A · Cloutier · 1999 [cited by applicant]
US 6161219A · Ramkumar et al. · 2000 [cited by applicant]
US 6212112B1 · Naura et al. · 2001 [cited by applicant]
US 6332186B1 · Elwood et al. · 2001 [cited by applicant]
US 6877020B1 · Bratt et al. · 2005 [cited by applicant]
US 7003542B2 · Devir · 2006 [cited by applicant]
US 7209939B2 · Castrapel et al. · 2007 [cited by applicant]
US 7725521B2 · Chen et al. · 2010 [cited by applicant]
US 7792895B1 · Juffa et al. · 2010 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 7912889B1 · Juffa et al. · 2011 [cited by applicant]
US 7932910B2 · Hansen et al. · 2011 [cited by applicant]
US 8392487B1 · Mesh et al. · 2013 [cited by applicant]
US 8984043B2 · Ginzburg et al. · 2015 [cited by applicant]
US 9442723B2 · Yang et al. · 2016 [cited by applicant]
US 9906359B2 · Gueron · 2018 [cited by applicant]
US 9960907B2 · Gueron · 2018 [cited by applicant]
US 10535114B2 · Bolz · 2020 [cited by applicant]
US 20030126176A1 · Devir · 2003 [cited by applicant]
US 20040111587A1 · Nair et al. · 2004 [cited by applicant]
US 20050193050A1 · Sazegari · 2005 [cited by applicant]
US 20060101245A1 · Nair et al. · 2006 [cited by applicant]
US 20060190517A1 · Guerrero · 2006 [cited by applicant]
US 20070186210A1 · Hussain et al. · 2007 [cited by applicant]
US 20080071851A1 · Zohar et al. · 2008 [cited by applicant]
US 20080140994A1 · Khailany et al. · 2008 [cited by applicant]
US 20080208942A1 · Won et al. · 2008 [cited by applicant]
US 20090043836A1 · Dupaquis et al. · 2009 [cited by applicant]
US 20090292758A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300091A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300249A1 · Moyer et al. · 2009 [cited by applicant]
US 20100180100A1 · Lu et al. · 2010 [cited by applicant]
US 20100325187A1 · Juffa et al. · 2010 [cited by applicant]
US 20120011348A1 · Eichenberger · 2012 [cited by examiner]
US 20120079252A1 · Sprangle · 2012 [cited by applicant]
US 20120113133A1 · Shpigelblat · 2012 [cited by applicant]
US 20120137074A1 · Kim et al. · 2012 [cited by applicant]
US 20120254588A1 · Adrian et al. · 2012 [cited by applicant]
US 20120314774A1 · Yang et al. · 2012 [cited by applicant]
US 20130305020A1 · Valentine et al. · 2013 [cited by applicant]
US 20140149480A1 · Catanzaro et al. · 2014 [cited by applicant]
US 20140208067A1 · Bradbury et al. · 2014 [cited by applicant]
US 20150067302A1 · Gueron · 2015 [cited by applicant]
US 20150199266A1 · Franchetti et al. · 2015 [cited by applicant]
US 20160224510A1 · Moudgill · 2016 [cited by examiner]
US 20180113708A1 · Corbal et al. · 2018 [cited by applicant]
US 20180157464A1 · Lutz · 2018 [cited by examiner]
US 20180321938A1 · Boswell et al. · 2018 [cited by applicant]
US 20190042193A1 · Pasca · 2019 [cited by examiner]
US 20190042544A1 · Kashyap · 2019 [cited by examiner]
US 20190065146A1 · Heddes · 2019 [cited by examiner]
US 20190171448A1 · Chen · 2019 [cited by examiner]
US 20190340486A1 · Mills · 2019 [cited by examiner]
US 20190354568A1 · Lindberg · 2019 [cited by examiner]
US 20200117450A1 · Mansell · 2020 [cited by examiner]
EP 2112591A2 · 2009 [cited by applicant]
KR 1020110079495A · 2011 [cited by applicant]
WO 2004053841A2 · 2004 [cited by applicant]
WO 2004111587A1 · 2004 [cited by applicant]
WO 2013101018A1 · 2013 [cited by applicant]
WO 2016003740A1 · 2016 [cited by applicant]
WO 2016105727A1 · 2016 [cited by applicant]
WO 2018125250A1 · 2018 [cited by applicant]
WO 2018174925A1 · 2018 [cited by applicant]
Intention to Grant, EP App. No. 21217772.9, Apr. 30, 2024, 6 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, Jan. 22, 2019, 5 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, Mar. 11, 2019, 2 pages. [cited by applicant]
Decision to grant a European patent, EP App. No. 19201841.4, Oct. 27, 2022, 2 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 19201841.4, May 27, 2020, 8 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21217772.9, Apr. 25, 2022, 7 pages. [cited by applicant]
Extended European Search Report and Search Opinion, EP App. No. 23200278.2, Jan. 11, 2024, 09 pages. [cited by applicant]
Intel, “Architecture Instruction Set Extensions Programming Reference”, 319433-024, Feb. 2016, pp. 5-119-5-126. [cited by applicant]
Intention to grant, EP App. No. 19201841.4, Jun. 7, 2022, 7 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/036038, Jan. 17, 2019, 14 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/040546, Oct. 3, 2019, 10 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/036038, Sep. 5, 2017 , 15 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040534, Jan. 3, 2018, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040536, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040537, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040540, Jan. 3, 2018, 14 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040546, Jan. 24, 2018, 15 pages. [cited by applicant]
Lahr, David L., “Distractions: Timing Matrix Multiplication in SciDB and Setting the No. of Worker Instances in SciDB and Running Matrix Multiplication Piecemeal”, Available Online at <http://dllahr.blogspot.com/2012/11… [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/201,442, May 4, 2018, 11 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/398,200, Jul. 28, 2020, 16 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/186,387, May 6, 2020, 10 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/201,442, Dec. 14, 2018, 5 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/186,387, Nov. 20, 2020, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/216,566, Nov. 17, 2022, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/190,761, Oct. 4, 2023, 8 pages. [cited by applicant]
Office Action, EP App. No. 19201841.4, Nov. 10, 2021, 4 pages. [cited by applicant]
Office Action, EP App. No. 21217772.9, Dec. 18, 2023, 06 pages. [cited by applicant]
Wikipedia, “bfloat16 floating-point format”, available online at <https://en.wikipedia.org/w/index.php?title=Bfloat16_floating-point_format&oldid=856298357>, Aug. 24, 2018, 3 pages. [cited by applicant]
“BFLOAT 16 data type in TensorFlow”, Slow learning AIGC, Available Online at <https://mp.weixin.qq.com/s/eHeewCAO0noD2d-dXWidMA?Poc_token=HKbIX2ej08FPi2dCmjJFssPBukplaGIvN63dPFzN>, Sep. 24, 2018, 2 pages. [cited by applicant]
CN Office Action, including Search Report, Chinese App. No. 202210022539.1, Dec. 20, 2024, 18 pages (9 pages of English Translation and 9 pages of Original Document). [cited by applicant]
Office Action, EP App. No. 23200278.2, Nov. 22, 2024, 08 pages. [cited by applicant]