IP Library › Granted Patent US 11,714,642
Granted Patent B2
US 11,714,642 · App. 17/706,428 · Granted Aug 1, 2023

Systems, methods, and apparatuses for tile store

Inventors: Robert Valentine (Kiryat Tivon, IL); Menachem Adelman (Haifa, IL); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Bret L. Toll (Hillsboro, OR); Milind B. Girkar (Sunnyvale, CA); Zeev Sperber (Zichron Yackov, IL); Mark J. Charney (Lexington, MA); Rinat Rappoport (Haifa, IL); Jesus Corbal (King City, OR); Stanislav Shwartsman (Haifa, IL); Igor Yanover (Yokneam Illit, IL); Alexander F. Heinecke (San Jose, CA); Barukh Ziv (Haifa, IL); Dan Baum (Haifa, IL); Yuri Gebil (Nahariya, IL)
Assignee: Intel Corporation
G06F9/30036G06F7/485G06F7/4876G06F7/762G06F9/3001G06F9/3016G06F9/30032G06F9/30043G06F9/30109G06F9/30112G06F9/30134G06F9/30145G06F9/30149G06F9/30185G06F9/30196G06F9/3818G06F9/3836G06F17/16G06F2212/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,642
App. No.
17/706,428
Granted
Aug 1, 2023
Kind
B2
Abstract

Embodiments detailed herein relate to matrix operations. In particular, the loading of a matrix (tile) from memory. For example, support for a loading instruction is described in at least a form of decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and destination memory information, and execution circuitry to execute the decoded instruction to store each data element of configured rows of the identified source matrix operand to memory based on the destination memory information.

Claims (25)

1. An apparatus comprising:

decode circuitry to decode a single instruction having fields for an opcode to indicate execution circuitry is to save a context state to memory, wherein the context state is to include multidimensional matrix data of tiles according to two configuration bits, a first configuration bit to correspond to configuration data loaded in a tile configuration and the second configuration bit to correspond to matrix data; and

execution circuitry to execute the decoded single instruction to store the context state to memory.

2. The apparatus of claim 1 , wherein the two configuration bits are located in a control register.

3. The apparatus of claim 1 , wherein the execution circuitry is a part of an accelerator.

4. The apparatus of claim 1 , wherein the execution circuitry is a part of a processor.

5. The apparatus of claim 1 , wherein the execution circuitry is further to write zeros beyond a specified number of rows of the matrix data of the tiles.

6. The apparatus of claim 1 , wherein the matrix data of the tiles is to include garbage data in areas that are not configured for use in tile operations.

7. The apparatus of claim 1 , wherein the tiles are a plurality of registers configured to represent a matrix.

8. A method comprising:

decoding a single instruction having fields for an opcode to indicate execution circuitry is to save a context state to memory, wherein the context state is to include multidimensional matrix data of tiles according to two configuration bits, a first configuration bit to correspond to configuration data loaded in a tile configuration and the second configuration bit to correspond to matrix data; and

executing the decoded single instruction to store the context state to memory.

9. The method of claim 8 , wherein the two configuration bits are located in a control register.

10. The method of claim 8 , wherein a size of each data element of the matrix data is a doubleword.

11. The method of claim 8 , wherein a size of each data element of the matrix data is a word.

12. The method of claim 8 , wherein the executing is further to write zeros beyond a specified number of rows of the matrix data of the tiles.

13. The method of claim 8 , wherein the matrix data of the tiles is to include garbage data in areas that are not configured for use in tile operations.

14. The method of claim 8 , wherein the tiles are a plurality of registers configured to represent a matrix.

15. A non-transitory machine-readable medium storing an instruction which causes an apparatus to perform a method, the method comprising:

decoding a single instruction having fields for an opcode to indicate execution circuitry is to save a context state to memory, wherein the context state is to include multidimensional matrix data of tiles according to two configuration bits, a first configuration bit to correspond to configuration data loaded in a tile configuration and the second configuration bit to correspond to matrix data; and

executing the decoded single instruction to store the context state to memory.

16. The non-transitory machine-readable medium of claim 15 , wherein the two configuration bits are located in a control register.

17. The non-transitory machine-readable medium of claim 15 , wherein the executing is further to write zeros beyond a specified number of rows of the matrix data of the tiles.

18. The non-transitory machine-readable medium of claim 15 , wherein the matrix data of the tiles is to include garbage data in areas that are not configured for use in tile operations.

19. The non-transitory machine-readable medium of claim 15 , wherein the tiles are a plurality of registers configured to represent a matrix.

Continuity (3)
Continuation 16487755
Provisional Application 62473732 · Mar 20, 2017
Related Publication 20220291927A1 · Sep 15, 2022
Cited By (5)
US 12,260,213 US 12,282,773 US 12,314,717 US 12,536,020 US 12,650,839