IP Library Granted Patent US 11,263,008
Granted Patent B2
US 11,263,008 · App. 16/487,774 · Granted Mar 1, 2022

Systems, methods, and apparatuses for tile broadcast

Inventors: Robert Valentine (Kiryat Tivon, IL); Zeev Sperber (Zichron Yackov, IL); Mark J. Charney (Lexington, MA); Bret L. Toll (Hillsboro, OR); Jesus Corbal (King City, OR); Alexander Heinecke (San Jose, CA); Barukh Ziv (Haifa, IL); Dan Baum (Haifa, IL); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Stanislav Shwartsman (Haifa, IL)
Assignee: Intel Corporation
G06F9/30036G06F7/485G06F7/4876G06F7/762G06F9/3001G06F9/3016G06F9/30043G06F9/30109G06F9/30112G06F9/30134G06F9/30145G06F9/30149G06F9/30185G06F9/30196G06F9/3818G06F9/3836G06F17/16G06F2212/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,263,008
App. No.
16/487,774
Granted
Mar 1, 2022
Kind
B2
Abstract

Embodiments detailed herein relate to matrix operations. In particular, embodiment of broadcasting elements are described. For example, some embodiments describe broadcasting a scalar to all configured data element positons of a destination matrix (tile). For example, some embodiments describe broadcasting a row to all configured data element positons of a destination matrix (tile). For example, some embodiments describe broadcasting a column to all configured data element positons of a destination matrix (tile).

Claims (31)

1. A processor comprising:

decode circuitry to decode an instruction having fields for an opcode, an identifier for a source operand, and an identifier for a destination matrix operand; and

execution circuitry to execute the decoded instruction according to the opcode to broadcast a single data element from the identified source operand into each configured data element of the identified destination matrix operand.

2. The processor of claim 1 , wherein the execution circuitry comprises a crossbar switch.

3. The processor of claim 1 , wherein the source operand is stored in memory.

4. The processor of claim 1 , wherein the source operand is stored in a packed data register.

5. The processor of claim 1 , wherein the source operand is stored in a general-purpose data register.

6. The processor of claim 1 , wherein the identified destination matrix operand is a plurality of registers to represent a two-dimensional matrix.

7. The processor of claim 1 , wherein the data elements of the row of data are doubleword in size.

8. The processor of claim 1 , wherein the data elements of the row of data are word in size.

9. The processor of claim 1 , further comprising storage to store a configuration of the identified destination matrix operand.

10. The processor of claim 1 , wherein the configuration is to indicate a number of rows and columns to be used by the identified destination matrix operand.

11. The processor of claim 1 , wherein the execution circuitry is to further zero unconfigured data elements of the identified destination matrix operand.

12. A method comprising:

decoding an instruction having fields for an opcode, an identifier for a source operand, and an identifier for a destination matrix operand; and

executing the decoded instruction according to the opcode to broadcast a single data element from the identified source operand into each configured data element of the identified destination matrix operand.

13. The method of claim 12 , wherein the source operand is stored in memory.

14. The method of claim 12 , wherein the source operand is stored in a register.

15. The method of claim 12 , wherein the identified destination matrix operand is a plurality of registers to represent a two-dimensional matrix.

16. The method of claim 12 , wherein the data elements of the row of data are doubleword in size.

17. The method of claim 12 , wherein the data elements of the row of data are word in size.

18. The method of claim 12 , further comprising:

zeroing unconfigured data elements of the identified destination matrix operand.

19. A non-transitory machine-readable medium storing an instruction which causes a processor to perform a method, the method comprising:

decoding the instruction having fields for an opcode, an identifier for a source operand, and an identifier for a destination matrix operand; and

executing the decoded instruction according to the opcode to broadcast a single data element from the identified source operand into each configured data element of the identified destination matrix operand.

20. A system comprising:

a processor;

an accelerator coupled to the processor, the accelerator including:

decode circuitry to decode an instruction having fields for an opcode, an identifier for a source operand, and an identifier for a destination matrix operand; and

execution circuitry to execute the decoded instruction according to the opcode to broadcast a single data element from the identified source operand into each configured data element of the identified destination matrix operand.

Continuity (2)
Provisional Application 62473732 · Mar 20, 2017
Related Publication 20200249947A1 · Aug 6, 2020
Cited By (5)
US 12,260,213 US 12,282,773 US 12,314,717 US 12,536,020 US 12,650,839