IP Library › Granted Patent US 10,338,919
Granted Patent B2
US 10,338,919 · App. 15/826,435 · Granted Jul 2, 2019

Generalized acceleration of matrix multiply accumulate operations

Inventors: Brent Ralph Boswell (Aloha, OR); Ming Y. Siu (Santa Clara, CA); Jack H. Choquette (Palo Alto, CA); Jonah M. Alben (San Jose, CA); Stuart Oberman (Sunnyvale, CA)
Assignee: NVIDIA Corporation
G06F9/30014G06F9/3001G06F9/3012G06F9/30036G06F9/3851G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,338,919
App. No.
15/826,435
Filed
Nov 29, 2017
Granted
Jul 2, 2019
Kind
B2
Examiner
MAI, TAN V
Art Unit
2182
USPC
708/607
Abstract

A method, computer readable medium, and processor are disclosed for performing matrix multiply and accumulate (MMA) operations. The processor includes a datapath configured to execute the MMA operation to generate a plurality of elements of a result matrix at an output of the datapath. Each element of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors associated with matrix operands specified in an instruction for the MMA operation. A dot product operation includes the steps of: generating a plurality of partial products by multiplying each element of a first vector with a corresponding element of a second vector; aligning the plurality of partial products based on the exponents associated with each element of the first vector and each element of the second vector; and accumulating the plurality of aligned partial products into a result queue utilizing at least one adder.

Claims (40)

1. A processor, comprising:

a datapath configured to execute a matrix multiply and accumulate (MMA) operation to generate a plurality of elements of a result matrix at an output of the datapath,

wherein each element in the plurality of elements of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors associated with matrix operands specified in an instruction for the MMA operation, and

wherein a dot product operation for calculating each dot product in the at least one dot product comprises:

generating a plurality of partial products by multiplying each element of a first vector with a corresponding element of a second vector,

aligning the plurality of partial products based on the exponents associated with each element of the first vector and each element of the second vector, and

accumulating the plurality of aligned partial products in parallel into a result queue utilizing at least one adder.

2. The processor of claim 1 , wherein the datapath is configured to convert one or more elements of the matrix operands to a half-precision, floating-point format value in a first stage of the datapath.

3. The processor of claim 1 , wherein the datapath is configured to generate a plurality of 4-vector dot products in parallel.

4. The processor of claim 1 , wherein the datapath includes a number of pipeline stages, and wherein at least one pipeline stage in the number of pipeline stages is shared with a double-precision, floating-point fused multiply accumulate (DFMA) datapath.

5. The processor of claim 4 , wherein the at least one pipeline stage comprises a completion adder that is configured to accumulate the plurality of aligned partial products in parallel into the result queue.

6. The processor of claim 1 , wherein the datapath is configured to generate the plurality of elements of the result matrix in a plurality of passes during a single instruction cycle.

7. The processor of claim 1 , wherein the processor is a parallel processing unit comprising a plurality of streaming multi-processors (SMs), each SM in the plurality of SMs including a register file and a number of cores, each core in the number of cores including an instance of the datapath.

8. The processor of claim 7 , wherein the MMA operation is configured to be executed by a number of threads in parallel, each thread configured to generate a portion of the elements in the result matrix on a particular core using different combinations of the vectors of the matrix operands specified in the instruction.

9. A method, comprising:

receiving an instruction for a matrix multiply and accumulate (MMA) operation; and

executing, by a processor, the MMA operation to generate a plurality of elements of a result matrix at an output of a datapath,

wherein each element in the plurality of elements of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors associated with matrix operands specified in the instruction for the MMA operation,

wherein a dot product operation for calculating each dot product in the at least one dot product comprises:

generating a plurality of partial products by multiplying each element of a first vector with a corresponding element of a second vector,

aligning the plurality of partial products based on the exponents associated with each element of the first vector and each element of the second vector, and

accumulating the plurality of aligned partial products into a result queue utilizing at least one adder.

10. The method of claim 9 , wherein the datapath is configured to convert one or more elements of the matrix operands to a half-precision, floating-point format value in a first stage of the datapath.

11. The method of claim 9 , wherein the datapath is configured to generate a plurality of 4-vector dot products in parallel.

12. The method of claim 9 , wherein the datapath includes a number of pipeline stages, and wherein at least one pipeline stage in the number of pipeline stages is shared with a double-precision, floating-point fused multiply accumulate (DFMA) datapath.

13. The method of claim 9 , wherein the datapath is configured to generate the plurality of elements of the result matrix in a plurality of passes during a single instruction cycle.

14. The method of claim 9 , wherein the processor is a parallel processing unit comprising a plurality of streaming multi-processors (SMs), each SM in the plurality of SMs including a register file and a number of cores, each core in the number of cores including an instance of the datapath.

15. The method of claim 14 , wherein the MMA operation is configured to be executed by a number of threads in parallel, each thread configured to generate a portion of the elements in the result matrix on a particular core using different combinations of the vectors of the matrix operands specified in the instruction.

16. A non-transitory, computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform steps comprising:

receiving an instruction for a matrix multiply and accumulate (MMA) operation; and

executing the MMA operation to generate a plurality of elements of a result matrix at an output of a datapath,

wherein each element in the plurality of elements of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors associated with matrix operands specified in the instruction for the MMA operation,

wherein a dot product operation for calculating each dot product in the at least one dot product comprises:

generating a plurality of partial products by multiplying each element of a first vector with a corresponding element of a second vector,

aligning the plurality of partial products based on the exponents associated with each element of the first vector and each element of the second vector, and

accumulating the plurality of aligned partial products into a result queue utilizing at least one adder.

17. The non-transitory, computer-readable storage medium of claim 16 , wherein the datapath is configured to convert one or more elements of the matrix operands to a half-precision, floating-point format value in a first stage of the datapath.

18. The non-transitory, computer-readable storage medium of claim 16 , wherein the datapath includes a number of pipeline stages, and wherein at least one pipeline stage in the number of pipeline stages is shared with a double-precision, floating-point fused multiply accumulate (DFMA) datapath.

19. The non-transitory, computer-readable storage medium of claim 16 , wherein the processor is a parallel processing unit comprising a plurality of streaming multi-processors (SMs), each SM in the plurality of SMs including a register file and a number of cores, each core in the number of cores including an instance of the datapath.

20. The non-transitory, computer-readable storage medium of claim 19 , wherein the MMA operation is configured to be executed by a number of threads in parallel, each thread configured to generate a portion of the elements in the result matrix on a particular core using different combinations of the vectors of the matrix operands specified in the instruction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2018
From: BOSWELL, BRENT RALPH; SIU, MING Y.; CHOQUETTE, JACK H.; ALBEN, JONAH M.; OBERMAN, STUART
To: NVIDIA CORPORATION
Reel/Frame 044539/0788 →
Continuity (2)
Provisional Application 62503159 · May 8, 2017
Related Publication 20180321938A1 · Nov 8, 2018
Cited By (8)
US 12,204,897 US 12,321,743 US 12,670,121 US 12,688,146 US 12,710,961 US 12,730,759 US 12,737,317 US 12,737,318