IP Library › Granted Patent US 12,405,770
Granted Patent B2
US 12,405,770 · App. 18/607,024 · Granted Sep 2, 2025

Matrix transpose and multiply

Inventors: Menachem Adelman (Modi'in, IL); Robert Valentine (Kiryat Tivon, IL); Barukh Ziv (Haifa, IL); Amit Gradstein (Binyamina, IL); Simon Rubanovich (Haifa, IL); Zeev Sperber (Zichron Yackov, IL); Mark J. Charney (Lexington, MA); Christopher J. Hughes (Santa Clara, CA); Alexander F. Heinecke (San Jose, CA); Evangelos Georganas (San Jose, CA); Binh Pham (Burlingame, CA)
Assignee: Intel Corporation
G06F7/78G06F9/3001G06F9/30036G06F9/30038G06F9/3016G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,770
App. No.
18/607,024
Granted
Sep 2, 2025
Kind
B2
Abstract

Embodiments for a matrix transpose and multiply operation are disclosed. In an embodiment, a processor includes a decoder and execution circuitry. The decoder is to decode an instruction having a format including an opcode field to specify an opcode, a first destination operand field to specify a destination matrix location, a first source operand field to specify a first source matrix location, and a second source operand field to specify a second source matrix location. The execution circuitry is to, in response to the decoded instruction, transpose the first source matrix to generate a transposed first source matrix, perform a matrix multiplication using the transposed first source matrix and the second source matrix to generate a result, and store the result in a destination matrix location.

Claims (37)

1. A processor, comprising:

a plurality of registers to store a plurality of packed data elements including a first plurality of packed data elements of a first source matrix tile and a second plurality of packed data elements of a second source matrix tile, the first and second source matrix tiles comprising respective portions of a first source matrix and a second source matrix, and wherein each packed data element of the plurality of packed data elements has an element width;

a decoder to decode one or more instructions, at least one instruction of the one or more instructions including an opcode field configured to specify an opcode, a first source operand configured to indicate the first plurality of packed data elements, a second source operand configured to indicate the second plurality of packed data elements, and a destination operand configured to indicate a result matrix tile;

execution circuitry, in response to the one or more instructions, to transpose the first source matrix tile in accordance with a granularity equal to the element width to generate a first transposed source matrix tile comprising the first plurality of packed data elements and to multiply the first transposed source matrix tile and the second source matrix tile, the execution circuitry comprising:

a plurality of multipliers to multiply data elements of the first transposed source matrix tile and corresponding data elements of the second source matrix tile to produce a corresponding plurality of products; and

one or more accumulators to add groups of the products to generate corresponding result data elements in the result matrix tile.

2. The processor of claim 1 , wherein to transpose the first source matrix tile, the execution circuitry is to transform each row of the first source matrix tile to a column of the first transposed source matrix tile.

3. The processor of claim 2 , wherein each column of the first transposed source matrix tile has a width equal to the element width.

4. The processor of claim 3 , wherein the element width comprises an 8-bit width or a 16-bit width and a result element width of each of the result data elements comprises a 32-bit width.

5. The processor of claim 4 , wherein the first and second plurality of packed data elements comprise 16-bit floating point data elements and each of the result data elements comprise single-precision floating point data elements.

6. The processor of claim 5 , wherein the 16-bit floating point data elements comprise bfloat16 data elements.

7. The processor of claim 1 , further comprising:

a scheduler to schedule the one or more instructions for execution on the execution circuitry.

8. A method comprising:

storing in a plurality of registers a plurality of packed data elements including a first plurality of packed data elements of a first source matrix tile and a second plurality of packed data elements of a second source matrix tile, the first and second source matrix tiles comprising respective portions of a first source matrix and a second source matrix, and wherein each packed data element of the plurality of packed data elements has an element width;

decoding one or more instructions, at least one instruction of the one or more instructions including an opcode field configured to specify an opcode, a first source operand configured to indicate the first plurality of packed data elements, a second source operand configured to indicate the second plurality of packed data elements, and a destination operand configured to indicate a result matrix tile;

transposing, in response to the one or more instructions, the first source matrix tile in accordance with a granularity equal to the element width to generate a first transposed source matrix tile comprising the first plurality of packed data elements and to multiply the first transposed source matrix tile and the second source matrix tile, wherein multiplying the first transposed source matrix tile and the second source matrix tile comprises:

multiplying data elements of the first transposed source matrix tile and corresponding data elements of the second source matrix tile to produce a corresponding plurality of products; and

adding groups of the products to generate corresponding result data elements in the result matrix tile.

9. The method of claim 8 , wherein transposing the first source matrix tile comprises transforming each row of the first source matrix tile to a column of the first transposed source matrix tile.

10. The method of claim 9 , wherein each column of the first transposed source matrix tile has a width equal to the element width.

11. The method of claim 10 , wherein the element width comprises an 8-bit width or a 16-bit width and a result element width of each of the result data elements comprises a 32-bit width.

12. The method of claim 11 , wherein the first and second plurality of packed data elements comprise 16-bit floating point data elements and each of the result data elements comprise single-precision floating point data elements.

13. The method of claim 12 , wherein the 16-bit floating point data elements comprise bfloat16 data elements.

14. The method of claim 8 , further comprising:

scheduling the one or more instructions for execution on the execution circuitry.

15. A non-transitory machine-readable medium containing instructions which, when executed by a processor, are to cause the processor to respond by:

storing in a plurality of registers a plurality of packed data elements including a first plurality of packed data elements of a first source matrix tile and a second plurality of packed data elements of a second source matrix tile, the first and second source matrix tiles comprising respective portions of a first source matrix and a second source matrix, and wherein each packed data element of the plurality of packed data elements has an element width;

decoding one or more instructions, at least one instruction of the one or more instructions including an opcode field configured to specify an opcode, a first source operand configured to indicate the first plurality of packed data elements, a second source operand configured to indicate the second plurality of packed data elements, and a destination operand configured to indicate a result matrix tile;

transposing, in response to the one or more instructions, the first source matrix tile in accordance with a granularity equal to the element width to generate a first transposed source matrix tile comprising the first plurality of packed data elements and to multiply the first transposed source matrix tile and the second source matrix tile, wherein multiplying the first transposed source matrix tile and the second source matrix tile comprises:

multiplying data elements of the first transposed source matrix tile and corresponding data elements of the second source matrix tile to produce a corresponding plurality of products; and

adding groups of the products to generate corresponding result data elements in the result matrix tile.

16. The machine-readable medium of claim 15 , wherein transposing the first source matrix tile comprises transforming each row of the first source matrix tile to a column of the first transposed source matrix tile.

17. The machine-readable medium of claim 16 , wherein each column of the first transposed source matrix tile has a width equal to the element width.

18. The machine-readable medium of claim 17 , wherein the element width comprises an 8-bit width or a 16-bit width and a result element width of each of the result data elements comprises a 32-bit width.

19. The machine-readable medium of claim 18 , wherein the first and second plurality of packed data elements comprise 16-bit floating point data elements and each of the result data elements comprise single-precision floating point data elements.

20. The machine-readable medium of claim 19 , wherein the 16-bit floating point data elements comprise bfloat16 data elements.

Continuity (2)
Continuation 16914318 · Jun 27, 2020
Related Publication 20240329938A1 · Oct 3, 2024
References Cited (82)
US 5247632A · Newman · 1993 [cited by applicant]
US 5475822A · Sibigtroth et al. · 1995 [cited by applicant]
US 5892962A · Cloutier · 1999 [cited by applicant]
US 6161219A · Ramkumar et al. · 2000 [cited by applicant]
US 6212112B1 · Naura et al. · 2001 [cited by applicant]
US 6332186B1 · Elwood et al. · 2001 [cited by applicant]
US 6877020B1 · Bratt et al. · 2005 [cited by applicant]
US 7003542B2 · Devir · 2006 [cited by applicant]
US 7209939B2 · Castrapel et al. · 2007 [cited by applicant]
US 7725521B2 · Chen et al. · 2010 [cited by applicant]
US 7792895B1 · Juffa et al. · 2010 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 7912889B1 · Juffa et al. · 2011 [cited by applicant]
US 7932910B2 · Hansen et al. · 2011 [cited by applicant]
US 8392487B1 · Mesh et al. · 2013 [cited by applicant]
US 8984043B2 · Ginzburg et al. · 2015 [cited by applicant]
US 9442723B2 · Yang et al. · 2016 [cited by applicant]
US 9906359B2 · Gueron · 2018 [cited by applicant]
US 9960907B2 · Gueron · 2018 [cited by applicant]
US 10535114B2 · Bolz · 2020 [cited by applicant]
US 11972230B2 · Adelman · 2024 [cited by examiner]
US 20030126176A1 · Devir · 2003 [cited by applicant]
US 20040111587A1 · Nair et al. · 2004 [cited by applicant]
US 20050193050A1 · Sazegari · 2005 [cited by applicant]
US 20060101245A1 · Nair et al. · 2006 [cited by applicant]
US 20060190517A1 · Guerrero · 2006 [cited by applicant]
US 20070186210A1 · Hussain et al. · 2007 [cited by applicant]
US 20080071851A1 · Zohar et al. · 2008 [cited by applicant]
US 20080140994A1 · Khailany et al. · 2008 [cited by applicant]
US 20080208942A1 · Won et al. · 2008 [cited by applicant]
US 20090043836A1 · Dupaquis et al. · 2009 [cited by applicant]
US 20090248778A1 · Magerlein · 2009 [cited by applicant]
US 20090292758A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300091A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300249A1 · Moyer et al. · 2009 [cited by applicant]
US 20100180100A1 · Lu et al. · 2010 [cited by applicant]
US 20100325187A1 · Juffa et al. · 2010 [cited by applicant]
US 20120079252A1 · Sprangle · 2012 [cited by applicant]
US 20120113133A1 · Shpigelblat · 2012 [cited by applicant]
US 20120137074A1 · Kim et al. · 2012 [cited by applicant]
US 20120254588A1 · Adrian et al. · 2012 [cited by applicant]
US 20120314774A1 · Yang et al. · 2012 [cited by applicant]
US 20130305020A1 · Valentine et al. · 2013 [cited by applicant]
US 20140149480A1 · Catanzaro et al. · 2014 [cited by applicant]
US 20150067302A1 · Gueron · 2015 [cited by applicant]
US 20150199266A1 · Franchetti et al. · 2015 [cited by applicant]
US 20150378734A1 · Hansen et al. · 2015 [cited by applicant]
US 20180113708A1 · Corbal et al. · 2018 [cited by applicant]
US 20190042202A1 · Sade et al. · 2019 [cited by applicant]
US 20190042254A1 · Sade et al. · 2019 [cited by applicant]
US 20190377573A1 · Magklis et al. · 2019 [cited by applicant]
CN 110727911A · 2020 [cited by applicant]
KR 1020110079495A · 2011 [cited by applicant]
WO 2004053841A2 · 2004 [cited by applicant]
WO 2016003740A1 · 2016 [cited by applicant]
WO 2016105727A1 · 2016 [cited by applicant]
WO 2018125250A1 · 2018 [cited by applicant]
WO 2018126073A1 · 2018 [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, filed Jan. 22, 2019, 5 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, filed Mar. 11, 2019, 2 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 20209948.7, May 31, 2021, 08 pages. [cited by applicant]
Intention to Grant, EP App. No. 20209948.7, May 2, 2024, 6 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/036038, Jan. 17, 2019, 14 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/040546, Oct. 3, 2019, 10 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/036038, Sep. 5, 2017 , 15 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040534, Jan. 3, 2018, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040536, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040537, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040540, Jan. 3, 2018, 14 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040546, Jan. 24, 2018, 15 pages. [cited by applicant]
Lahr, David L., “Distractions: Timing Matrix Multiplication in SciDB and Setting the Number of Worker Instances in SciDB and Running Matrix Multiplication Piecemeal”, Available Online at <http://dllahr.blogspot.com/2012… [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/201,442, May 4, 2018, 11 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/398,200, Jul. 28, 2020, 17 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/914,318, Aug. 16, 2023, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/201,442, Dec. 14, 2018, 5 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/914,318, Dec. 22, 2023, 8 pages. [cited by applicant]
Office Action, EP App. No. 20209948.7, Apr. 17, 2023, 9 pages. [cited by applicant]
Office Action, EP App. No. 20209948.7, Jun. 14, 2022, 5 pages. [cited by applicant]
Summons to attend oral proceedings, EP App. No. 20209948.7, Sep. 27, 2023, 20 pages. [cited by applicant]
Decision to Grant, EP App. No. 20209948.7, Sep. 12, 2024, 2 pages. [cited by applicant]
Extended European search report and Search Opinion , EP App. No. 24203555.8, Jan. 21, 2025, 10 pages. [cited by applicant]
Extended European search report and Search Opinion , EP App. No. 24205150.6, Jan. 21, 2025, 36 pages. [cited by applicant]