IP Library › Granted Patent US 10,769,238
Granted Patent B2
US 10,769,238 · App. 16/576,144 · Granted Sep 8, 2020

Matrix multiplication on a systolic array

Inventors: Chia-Yu Chen (White Plains, NY); Jungwook Choi (Elmsford, NY); Kailash Gopalakrishnan (San Jose, CA); Victor Han (San Diego, CA); Vijayalakshmi Srinivasan (New York, NY); Jintao Zhang (Princeton, NJ)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,769,238
App. No.
16/576,144
Granted
Sep 8, 2020
Kind
B2
Abstract

Techniques facilitating matrix multiplication on a systolic array are provided. A computer-implemented method can comprise populating, by a system operatively coupled to a processor, respective first registers of one or more processing elements of a systolic array structure with respective input data bits of a first data matrix. The one or more processing elements can comprise a first processing element that comprises a first input data bit of the first data matrix and a first activation bit of a second data matrix. The method can also include determining, by the system, at the first processing element, a first partial sum of a third data matrix. Further, the method can include streaming, by the system, the first partial sum of the third data matrix from the first processing element.

Claims (36)

1. A system, comprising:

a memory that stores computer executable components; and

a processor that executes the computer executable components stored in the memory, wherein the computer executable components comprise:

a load manager component that populates respective first registers of all processing elements of a systolic array structure with respective input data bits of a first data matrix, wherein the load manager component further inputs a first activation bit of a second data matrix into a first processing element of the processing elements, and the respective input data bits of the first data matrix are maintained in the respective first registers while a matrix multiplication of the first data matrix and the second data matrix is completed.

2. The system of claim 1 , further comprising:

a computation component that determines, during the matrix multiplication, at the first processing element, a first partial sum of a third data matrix based on a first product of the first activation bit and a first input data bit of the first data matrix, and a first initial value of the third data matrix.

3. A computer-implemented method, comprising:

populating, by a system operatively coupled to a processor, respective first registers of all processing elements of a systolic array structure with respective input data bits of a first data matrix;

inputting, by the system, a first activation bit of a second data matrix into a first processing element of the processing elements; and

maintaining, by the system, respective input data bits of the first data matrix in the respective first registers while a matrix multiplication of the first data matrix and the second data matrix is completed.

4. The computer-implemented method of claim 3 , further comprising:

determining, by the system during the matrix multiplication, at the first processing element, a first partial sum of a third data matrix based on a first product of the first activation bit and a first input data bit of the first data matrix, and a first initial value of the third data matrix.

5. The computer-implemented method of claim 4 , further comprising:

streaming, by the system during the matrix multiplication, the first partial sum of the third data matrix along a first dimension to a second processing element of the processing elements.

6. The computer-implemented method of claim 5 , further comprising:

determining, by the system during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of the first partial sum and a second product determined based on a second activation bit of the second data matrix and a second input data bit of the first data matrix stored in the first register of the second processing element.

7. The computer-implemented method of claim 6 , further comprising:

streaming, by the system during the matrix multiplication, the second partial sum of the third data matrix from the second processing element and along the first dimension.

8. The computer-implemented method of claim 5 , further comprising:

determining, by the system during the matrix multiplication, during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of a second product and a second initial value of the third data matrix, wherein the second product is determined based on the first activation bit and a second input data bit of the first data matrix stored in the first register of the second processing element.

9. The computer-implemented method of claim 8 , further comprising: streaming, by the system during the matrix multiplication, the second partial sum of the third data matrix from the second processing element and along the first dimension.

10. A computer program product for facilitating matrix multiplication on a systolic array structure, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processing component to cause the processing component to:

populate respective first registers of all processing elements of a systolic array structure with respective input data bits of a first data matrix;

input a first activation bit of a second data matrix into a first processing element of the processing elements; and

maintain respective input data bits of the first data matrix in the respective first registers while a matrix multiplication of the first data matrix and the second data matrix is completed.

11. The computer program product of claim 10 , wherein the program instructions further cause the processing component to:

determine, during the matrix multiplication, at the first processing element, a first partial sum of a third data matrix based on a first product of the first activation bit and a first input data bit of the first data matrix, and a first initial value of the third data matrix.

12. The computer program product of claim 11 , wherein the program instructions further cause the processing component to:

stream, during the matrix multiplication, the first partial sum of the third data matrix along a first dimension to a second processing element of the processing elements.

13. The computer program product of claim 12 , wherein the program instructions further cause the processing component to:

determine, during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of the first partial sum and a second product determined based on a second activation bit of the second data matrix and a second input data bit of the first data matrix stored in the first register of the second processing element.

14. The computer program product of claim 13 , wherein the program instructions further cause the processing component to:

stream, during the matrix multiplication, the second partial sum of the third data matrix from the second processing element and along the first dimension.

15. The computer program product of claim 12 , wherein the program instructions further cause the processing component to:

determine, during the matrix multiplication, during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of a second product and a second initial value of the third data matrix, wherein the second product is determined based on the first activation bit and a second input data bit of the first data matrix stored in the first register of the second processing element; and

stream, during the matrix multiplication, the second partial sum of the third data matrix from the second processing element and along the first dimension.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2019
From: CHEN, CHIA-YU; CHOI, JUNGWOOK; GOPALAKRISHNAN, KAILASH; HAN, VICTOR; SRINIVASAN, VIJAYALAKSHMI; ZHANG, JINTAO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050442/0358 →
Continuity (4)
Continuation 16381530 · Apr 11, 2019
Continuation 15842422 · Dec 14, 2017
Continuation 15460755 · Mar 16, 2017
Related Publication 20200012706A1 · Jan 9, 2020
Cited By (5)
US 12,423,058 US 12,517,700 US 12,645,425 US 12,663,966 US 12,681,693