IP Library Granted Patent US 12,353,878
Granted Patent B2
US 12,353,878 · App. 17/359,519 · Granted Jul 8, 2025

Apparatuses, methods, and systems for instructions for matrix multiplication instructions

Inventors: Menachem Adelman (Haifa, IL); Robert Valentine (Kiryat Tivon, IL); Zeev Sperber (Zikhron Yaakov, IL); Amit Gradstein (Binyamina, IL); Simon Rubanovich (Haifa, IL); Sagi Meller (Zikhron Yaakov, IL); Christopher Hughes (Santa Clara, CA); Evangelos Georganas (San Mateo, CA); Alexander Heinecke (San Jose, CA); Mark Charney (Lexington, MA)
Assignee: Intel Corporation
G06F9/30036G06F9/3001G06F9/30025G06F9/30038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,878
App. No.
17/359,519
Granted
Jul 8, 2025
Kind
B2
Abstract

Techniques for matrix multiplication are described. In some examples, decode circuitry is to decode a single instruction having fields for an opcode, an indication of a location of a first source operand, an indication of a location of a second source operand, and an indication of a location of a destination operand, wherein the opcode is to indicate that execution circuitry is to at least convert data elements of the first and second source operands from a first floating point representation to a second floating point representation, perform matrix multiplication with the converted data elements, and accumulate results of the matrix multiplication in the destination operand in the first floating point representation; and the execution circuitry is to execute to the decoded instruction as specified by the opcode.

Claims (23)

1. An apparatus comprising:

decode circuitry to decode a single instruction having fields for an opcode, an indication of a location of a first source operand, an indication of a location of a second source operand, and an indication of a location of a destination operand, wherein the opcode is to indicate that execution circuitry is to at least convert data elements of the first and second source operands from a first floating point representation to a second floating point representation, perform matrix multiplication with the converted data elements, and accumulate results of the matrix multiplication in the destination operand in the first floating point representation, wherein the first floating point representation is single precision floating point and the second floating point representation is a 19-bit floating-point representation; and

the execution circuitry to execute to the decoded instruction as specified by the opcode.

2. The apparatus of claim 1 , wherein the execution circuitry is further to execute the decoded instruction to convert the data elements from the first and second source operands from a signaling not-a-number value to a quiet not-a-number value.

3. The apparatus of claim 1 , wherein the 19-bit floating-point representation has a sign bit, 8 exponent bits, and 10 mantissa bits that are corresponding most significant bits of a single precision floating number.

4. The apparatus of claim 1 , wherein the execution circuitry is further to execute the decoded instruction to zero any remaining data elements of the destination operand.

5. A non-transitory machine readable medium storing at least an instance of a single instruction, wherein a processor is to respond to the single instruction to perform a method comprising:

decoding the single instruction having fields for an opcode, an indication of a location of a first source operand, an indication of a location of a second source operand, and an indication of a location of a destination operand, wherein the opcode is to indicate that execution circuitry is to at least convert data elements of the first and second source operands from a first floating point representation to a second floating point representation, perform matrix multiplication with the converted data elements, and accumulate results of the matrix multiplication in the destination operand in the first floating point representation, wherein the first floating point representation is single precision floating point and the second floating point representation is a 19-bit floating point representation;

executing the decoded instruction to perform operations as specified by the opcode.

6. The non-transitory machine readable medium of claim 5 , wherein the executing further comprises converting the data elements from the first and second source operands from a signaling not-a-number value to a quiet not-a-number value.

7. The non-transitory machine readable medium of claim 5 , wherein the 19-bit floating-point representation has a lower 13 bits set to zero.

8. The non-transitory machine readable medium of claim 5 , wherein the executing further comprises zeroing any remaining data elements of the destination operand.

9. The non-transitory machine readable medium of claim 5 , further comprising:

translating the single instruction into one or more instructions of a different instruction set, wherein the executing the decoded instruction as specified by the opcode comprises executing the one or more instructions of the different instruction set to perform the operations as specified by the opcode of the single instruction.

10. An apparatus comprising:

decode circuitry to decode a single instruction having fields for an opcode, an indication of a location of a first source operand, an indication of a location of a second source operand, and an indication of a location of a destination operand, wherein the opcode is to indicate that execution circuitry is to at least convert data elements of the first and second source operands from a first floating point representation to a second floating point representation, perform matrix multiplication with a transpose of the converted data elements of the first source operand and non-transposed, converted data elements of the second source operand, and accumulate results of the matrix multiplication in the destination operand in the first floating point representation, wherein the first floating point representation is single precision floating point and the second floating point representation is a 19-bit floating point representation; and

the execution circuitry to execute to the decoded instruction as specified by the opcode.

11. The apparatus of claim 10 , wherein the execution circuitry is further to execute the decoded instruction to convert the data elements from the first and second source operands from a signaling not-a-number value to a quiet not-a-number value.

12. The apparatus of claim 10 , wherein the 19-bit floating-point representation has a lower 13 bits set to zero.

13. The apparatus of claim 10 , wherein the execution circuitry is further to execute the decoded instruction to zero any remaining data elements of the destination operand.

14. A non-transitory machine-readable medium storing at least an instance of a single instruction, wherein a processor is to respond to the single instruction to perform a method comprising:

decoding the single instruction having fields for an opcode, an indication of a location of a first source operand, an indication of a location of a second source operand, and an indication of a location of a destination operand, wherein the opcode is to indicate that execution circuitry is to at least convert data elements of the first and second source operands from a first floating point representation to a second floating point representation, perform matrix multiplication with a transpose of the converted data elements of the first source operand and non-transposed, converted data elements of the second source operand, and accumulate results of the matrix multiplication in the destination operand in the first floating point representation, wherein the first floating point representation is single precision floating point and the second floating point representation is a 19-bit floating point representation;

executing the decoded instruction as specified by the opcode.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2021
From: ADELMAN, MENACHEM; VALENTINE, ROBERT; SPERBER, ZEEV; GRADSTEIN, AMIT; RUBANOVICH, SIMON; MELLER, SAGI; HUGHES, CHRISTOPHER; GEORGANAS, EVANGELOS; HEINECKE, ALEXANDER; CHARNEY, MARK
To: INTEL CORPORATION
Reel/Frame 058600/0656 →
Continuity (1)
Related Publication 20220414182A1 · Dec 29, 2022
References Cited (13)
US 10223114B1 · Madduri et al. · 2019 [cited by applicant]
US 11126428B2 · Henry · 2021 [cited by examiner]
US 11544057B2 · Henry · 2023 [cited by examiner]
US 20080040584A1 · Hansen · 2008 [cited by examiner]
US 20180315399A1 · Kaul · 2018 [cited by examiner]
US 20190079768A1 · Heinecke et al. · 2019 [cited by applicant]
US 20220197601A1 · Adelman · 2022 [cited by examiner]
WO 2008036944A1 · 2008 [cited by applicant]
Kharya, Paresh, “TensorFloat-32 in the A100 GPU Accelerates AI Training, HPC up to 20x: NVIDIA's Ampere architecture with TF32 speeds single-precision work, maintaining accuracy and using No. new code”; 4 pages, May 14,… [cited by applicant]
Office Action, EP App. No. 22181066.6, Apr. 3, 2024, 4 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22181066.6, Nov. 16, 2022, 7 pages. [cited by applicant]
Office Action, EP App. No. 22181066.6, Nov. 6, 2023, 4 pages. [cited by applicant]
Office Action, EP App. No. 22181066.6, Nov. 26, 2024, 4 pages. [cited by applicant]