IP Library › Granted Patent US 12,020,028
Granted Patent B2
US 12,020,028 · App. 17/134,373 · Granted Jun 25, 2024

Apparatuses, methods, and systems for 8-bit floating-point matrix dot product instructions

Inventors: Naveen Mellempudi (Bangalore, IN); Alexander F. Heinecke (San Jose, CA); Robert Valentine (Kiryat Tivon, IL); Mark J. Charney (Lexington, MA); Christopher J. Hughes (Santa Clara, CA); Evangelos Georganas (San Mateo, CA); Zeev Sperber (Zichron Yackov, IL); Amit Gradstein (Binyamina, IL); Simon Rubanovich (Binyamina, IL)
Assignee: Intel Corporation
G06F9/30036G06F7/49915G06F9/30196G06F9/3887
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,020,028
App. No.
17/134,373
Granted
Jun 25, 2024
Kind
B2
Abstract

Systems, methods, and apparatuses relating to 8-bit floating-point matrix dot product instructions are described. A processor embodiment includes fetch circuitry to fetch an instruction having fields to specify an opcode and locations of a destination matrix having single-precision elements, a first source matrix, and a second source matrix, the source matrices having elements that each comprise a quadruple of 8-bit floating-point values, the opcode to indicate execution circuitry is to cause, for each element of the first source matrix and corresponding element of the second source matrix, a conversion of the 8-bit floating-point values to single-precision values, a multiplication of different pairs of converted single-precision values to generate plurality of results, and an accumulation of the results with previous contents of a corresponding element of the destination matrix, decode circuitry to decode the fetched instruction, and the execution circuitry to respond to the decoded instruction as specified by the opcode.

Claims (39)

1. An apparatus comprising:

fetch circuitry to fetch a single instruction having fields to specify an opcode and locations of a M by N destination matrix having single-precision elements, an M by K first source matrix, and a K by N second source matrix, the source matrices having elements that each comprise a quadruple of 8-bit floating-point values, the opcode to indicate execution circuitry is to cause, for each element of the first source matrix and corresponding element of the second source matrix, a conversion of the 8-bit floating-point values to single-precision values, a multiplication of converted single-precision values from first values of the quadruples together to generate a first result, a multiplication of converted single-precision values from second values of the quadruples together to generate a second result, a multiplication of converted single-precision values from third values of the quadruples together to generate a third result, a multiplication of converted single-precision values from fourth values of the quadruples together to generate a fourth result, and an accumulation of the first, second, third, and fourth results with previous contents of a corresponding element of the destination matrix;

decode circuitry to decode the fetched instruction; and

the execution circuitry to respond to the decoded instruction as specified by the opcode.

2. The apparatus of claim 1 , wherein a format of the 8-bit floating-point values is specified by the opcode of the single instruction.

3. The apparatus of claim 1 , wherein M, N, and K are specified by the single instruction.

4. The apparatus of claim 1 , where the execution circuitry is to cause a matrix operations accelerator to perform at least the multiplications and the accumulation.

5. The apparatus of claim 4 , wherein M, N, and K are specified by a configuration of the matrix operations accelerator to be programmed by execution of a matrix accelerator configuration instruction before executing the single instruction.

6. The apparatus of claim 1 , wherein the execution circuitry is further to cause saturation of execution results, as necessary.

7. The apparatus of claim 1 , wherein the single instruction is further to specify a writemask comprising M×N bits, each bit to control whether to mask a corresponding element of the destination matrix.

8. The apparatus of claim 1 , wherein the execution circuitry is further to generate a fault when a fault condition occurs, the fault condition selectable from:

the destination matrix having a fewer number of rows than a number of rows of the first source matrix; and

the destination matrix having a fewer number of columns than a number of columns of the second source matrix.

9. A method comprising:

fetching, by fetch circuitry of a processor, a single instruction having fields to specify an opcode and locations of a M by N destination matrix having single-precision elements, an M by K first source matrix, and a K by N second source matrix, the source matrices having elements that each comprise a quadruple of 8-bit floating-point values, the opcode to indicate execution circuitry is to cause, for each element of the first source matrix and corresponding element of the second source matrix, a conversion of the 8-bit floating-point values to single-precision values, a multiplication of converted single-precision values from first values of the quadruples together to generate a first result, a multiplication of converted single-precision values from second values of the quadruples together to generate a second result, a multiplication of converted single-precision values from third values of the quadruples together to generate a third result, a multiplication of converted single-precision values from fourth values of the quadruples together to generate a fourth result, and an accumulation of the first, second, third, and fourth results with previous contents of a corresponding element of the destination matrix;

decoding, by decode circuitry of the processor, the fetched instruction into a decoded single instruction; and

executing, by the execution circuitry of the processor, the decoded single instruction according to the opcode.

10. The method of claim 9 , wherein a format of the 8-bit floating-point values is specified by the opcode of the single instruction.

11. The method of claim 9 , wherein M, N, and K are specified by the single instruction.

12. The method of claim 9 , where the execution circuitry causes a matrix operations accelerator to perform at least the multiplications and the accumulation.

13. The method of claim 12 , further comprising executing, by the execution circuitry of the processor before executing the single instruction, a matrix accelerator configuration instruction that programs a configuration of the matrix operations accelerator specifying M, N, and K.

14. The method of claim 9 , wherein the executing comprises saturating the execution results.

15. The method of claim 9 , wherein the single instruction further specifies a writemask comprising M×N bits, each bit controlling whether to mask a corresponding element of the destination matrix.

16. The method of claim 9 , wherein the executing generates a fault when a fault condition occurs, the fault condition selectable from:

the destination matrix having a fewer number of rows than a number of rows of the first source matrix; and

the destination matrix having a fewer number of columns than a number of columns of the second source matrix.

17. A non-transitory machine readable medium that stores program code that when executed by a machine causes the machine to perform a method comprising:

fetching, by fetch circuitry of a processor, a single instruction having fields to specify an opcode and locations of a M by N destination matrix having single-precision elements, an M by K first source matrix, and a K by N second source matrix, the source matrices having elements that each comprise a quadruple of 8-bit floating-point values, the opcode to indicate execution circuitry is to cause, for each element of the first source matrix and corresponding element of the second source matrix, a conversion of the 8-bit floating-point values to single-precision values, a multiplication of converted single-precision values from first values of the quadruples together to generate a first result, a multiplication of converted single-precision values from second values of the quadruples together to generate a second result, a multiplication of converted single-precision values from third values of the quadruples together to generate a third result, a multiplication of converted single-precision values from fourth values of the quadruples together to generate a fourth result, and an accumulation of the first, second, third, and fourth results with previous contents of a corresponding element of the destination matrix;

decoding, by decode circuitry of the processor, the fetched instruction into a decoded single instruction; and

executing, by the execution circuitry of the processor, the decoded single instruction according to the opcode.

18. The non-transitory machine readable medium of claim 17 , wherein a format of the 8-bit floating-point values is specified by the opcode of the single instruction.

19. The non-transitory machine readable medium of claim 17 , wherein M, N, and K are specified by the single instruction.

20. The non-transitory machine readable medium of claim 17 , where the executing comprises the execution circuitry causing a matrix operations accelerator to perform at least the multiplications and the accumulation.

21. The non-transitory machine readable medium of claim 20 , wherein the method further comprises executing, by the execution circuitry of the processor before executing the single instruction, a matrix accelerator configuration instruction that programs a configuration of the matrix operations accelerator specifying M, N, and K.

22. The non-transitory machine readable medium of claim 17 , wherein the executing comprises saturating the execution results.

23. The non-transitory machine readable medium of claim 17 , wherein the single instruction further specifies a writemask comprising M×N bits, each bit controlling whether to mask a corresponding element of the destination matrix.

24. The non-transitory machine readable medium of claim 17 , wherein the executing generates a fault when a fault condition occurs, the fault condition selectable from:

the destination matrix having a fewer number of rows than a number of rows of the first source matrix; and

the destination matrix having a fewer number of columns than a number of columns of the second source matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2021
From: MELLEMPUDI, NAVEEN; HEINECKE, ALEXANDER F.; VALENTINE, ROBERT; CHARNEY, MARK J.; HUGHES, CHRISTOPHER J.; GEORGANAS, EVANGELOS; SPERBER, ZEEV; GRADSTEIN, AMIT; RUBANOVICH, SIMON
To: INTEL CORPORATION
Reel/Frame 055218/0736 →
Continuity (1)
Related Publication 20220206801A1 · Jun 30, 2022
Cited By (5)
US 12,260,213 US 12,282,773 US 12,314,717 US 12,536,020 US 12,650,839