IP Library Granted Patent US 12,124,847
Granted Patent B2
US 12,124,847 · App. 16/474,475 · Granted Oct 22, 2024

Systems, methods, and apparatuses for tile transpose

Inventors: Robert Valentine (Kiryat Tivon, IL); Dan Baum (Haifa, IL); Zeev Sperber (Zichron Yaakov, IL); Jesus Corbal (Barcelona, ES); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Bret L Toll (Hillsboro, OR); Mark J. Charney (Lexington, MA); Barukh Ziv (Haifa, IL); Alexander Heinecke (San Jose, CA); Milind Girkar (Sunnyvale, CA); Menachem Adelman (Haifa, IL); Simon Rubanovich (Haifa, IL)
Assignee: Intel Corporation
G06F9/30036G06F7/485G06F7/4876G06F7/762G06F9/3001G06F9/30032G06F9/30043G06F9/30109G06F9/30112G06F9/30134G06F9/30145G06F9/30149G06F9/3016G06F9/30185G06F9/30196G06F9/3818G06F9/3836G06F17/16G06F2212/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,124,847
App. No.
16/474,475
Granted
Oct 22, 2024
Kind
B2
Abstract

Embodiments detailed herein relate to matrix operations. In particular, support for a matrix transpose instruction is detailed. In some embodiments, decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and a destination matrix operand identifier; and execution circuitry to execute the decoded instruction to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand are detailed.

Claims (25)

1. A processor comprising:

decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier of a single two-dimensional tile register in a matrix operations accelerator of the processor, and a destination matrix operand identifier; and

execution circuitry to execute the decoded instruction to cause the matrix operations accelerator to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand and zero any remaining columns of the identified destination matrix operand and unconfigured rows of the identified destination matrix operand.

2. The processor of claim 1 , wherein the opcode defines a size of each data element of the source and destination matrix operands.

3. The processor of claim 2 , wherein the size of each data element of the source and destination matrix operands is a doubleword.

4. The processor of claim 2 , wherein the size of each data element of the source and destination matrix operands is a word.

5. The processor of claim 1 , wherein the destination matrix operand is a single two-dimensional tile register in the matrix operations accelerator of the processor.

6. The processor of claim 1 , wherein the execution circuitry is to fault upon a determination of one of: the identified source operand has a different number of rows than a number of columns of the identified destination operand, the identified source operand has a different number of columns than a number of rows in the identified destination operand, and the identified source and destination matrix operands have different sized data elements.

7. A method comprising:

decoding an instruction having fields for an opcode, a source matrix operand identifier of a single two-dimensional tile register in a matrix operations accelerator, and a destination matrix operand identifier; and

executing the decoded instruction to cause the matrix operations accelerator to transpose data elements of the identified source matrix operand into transposed data element positions of the identified destination matrix operand and zero any remaining columns of the identified destination matrix operand and unconfigured rows of the identified destination matrix operand.

8. The method of claim 7 , wherein the opcode defines a size of each data element of the source and destination matrix operands.

9. The method of claim 8 , wherein the size of each data element of the source and destination matrix operands is a doubleword.

10. The method of claim 8 , wherein the size of each data element of the source and destination matrix operands is a word.

11. The method of claim 7 , wherein the destination matrix operand is a single two-dimensional tile register in the matrix operations accelerator.

12. The method of claim 7 , further comprising:

faulting upon a determination of one of: the identified source operand has a different number of rows than a number of columns in the identified destination operand, the identified source operand has a different number of columns than a number of rows in the identified destination operand, and the identified source and destination matrix operands have different sized data elements.

13. A non-transitory machine-readable medium storing an instruction which causes a processor to perform a method, the method comprising:

decoding an instruction having fields for an opcode, a source matrix operand identifier of a single two-dimensional tile register in a matrix operations accelerator, and a destination matrix operand identifier; and

executing the decoded instruction to cause the matrix operations accelerator to transpose data elements of the identified source matrix operand into transposed data element positions of the identified destination matrix operand and zero any remaining columns of the identified destination matrix operand and unconfigured rows of the identified destination matrix operand.

14. The non-transitory machine-readable medium of claim 13 , wherein the opcode defines a size of each data element of the source and destination matrix operands.

15. The non-transitory machine-readable medium of claim 14 , wherein the size of each data element of the source and destination matrix operands is one of a word and a doubleword.

16. The non-transitory machine-readable medium of claim 13 , wherein the destination matrix operand is a single two-dimensional tile register in the matrix operations accelerator.

17. The non-transitory machine-readable medium of claim 13 , further comprising:

faulting upon a determination of one of: the identified source operand has a different number of rows than a number of columns of the identified destination operand, the identified source operand has a different number of columns than a number of rows in the identified destination operand, and the identified source and destination matrix operands have different sized data elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2019
From: VALENTINE, ROBERT; BAUM, DAN; SPERBER, ZEEV; CORBAL, JESUS; OULD-AHMED-VALL, ELMOUSTAPHA; TOLL, BRET L.; CHARNEY, MARK J.; ZIV, BARUKH; HEINECKE, ALEXANDER; GIRKAR, MILIND; ADELMAN, MENACHEM; RUBANOVICH, SIMON
To: INTEL CORPORATION
Reel/Frame 049718/0534 →
Continuity (2)
Provisional Application 62473732 · Mar 20, 2017
Related Publication 20190347100A1 · Nov 14, 2019