IP Library › Granted Patent US 10,261,978
Granted Patent B2
US 10,261,978 · App. 15/842,422 · Granted Apr 16, 2019

Matrix multiplication on a systolic array

Inventors: Chia-Yu Chen (White Plains, NY); Jungwook Choi (Elmsford, NY); Kailash Gopalakrishnan (San Jose, CA); Victor Han (San Diego, CA); Vijayalakshmi Srinivasan (New York, NY); Jintao Zhang (Princeton, NJ)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,261,978
App. No.
15/842,422
Granted
Apr 16, 2019
Kind
B2
Abstract

Techniques facilitating matrix multiplication on a systolic array are provided. A computer-implemented method can comprise populating, by a system operatively coupled to a processor, respective first registers of one or more processing elements of a systolic array structure with respective input data bits of a first data matrix. The one or more processing elements can comprise a first processing element that comprises a first input data bit of the first data matrix and a first activation bit of a second data matrix. The method can also include determining, by the system, at the first processing element, a first partial sum of a third data matrix. Further, the method can include streaming, by the system, the first partial sum of the third data matrix from the first processing element.

Claims (19)

1. A computer-implemented method, comprising:

populating, by a system operatively coupled to a processor, respective first registers of all processing elements of a systolic array structure with respective input data bits of a first data matrix, wherein the processing elements comprise a first processing element that comprises a first input data bit of the input data bits of the first data matrix and a first activation bit of a second data matrix, and wherein the systolic array structure comprises a first dimension and a second dimension of the processing elements, and the respective input data bits of the first data matrix are maintained in the respective first registers while a matrix multiplication of the first data matrix and the second data matrix is completed;

determining, by the system during the matrix multiplication, at the first processing element, a first partial sum of a third data matrix based on a first sum that comprises a first product and a first initial value of the third data matrix, wherein the first product is determined based on the first activation bit and the first input data bit; and

streaming, by the system during the matrix multiplication, the first partial sum of the third data matrix from the first processing element and along the second dimension.

2. The computer-implemented method of claim 1 , wherein the streaming the first partial sum of the third data matrix along the second dimension of the processing elements of the systolic array structure comprises inputting the first partial sum to a second processing element of the processing elements of the systolic array structure, and wherein a second input data bit of the input data bits of the first data matrix is stored in a second register of a second processing element, and wherein the computer-implemented method further comprising:

determining, by the system during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of a second product and the first partial sum, wherein the second product is determined based on a second activation bit of the second data matrix and the second input data bit; and

streaming, by the system during the matrix multiplication, the second activation bit from the second processing element and along the first dimension, and the second partial sum of the third data matrix from the second processing element and along the second dimension.

3. The computer-implemented method of claim 2 , further comprises:

shifting, by the system during the matrix multiplication, at a first clock cycle the first activation bit from the first processing element to a third processing element along the first dimension, and the first partial sum from the first processing element to the second processing element along the second dimension; and

shifting, by the system during the matrix multiplication, at a second clock cycle, the first activation bit from the second processing element to a fourth processing element along the first dimension, the second activation bit from the second processing element to a fifth processing element along the first dimension, and the second partial sum from the second processing element to a sixth processing element along the second dimension.

4. The computer-implemented method of claim 1 , wherein the streaming the first activation bit along the first dimension of the processing elements comprises inputting the first activation bit to a second processing element of the processing elements of the systolic array structure, and wherein the second processing element comprises a second input data bit of the input data bits of the first data matrix, and wherein the computer-implemented method further comprising:

determining, by the system during the matrix multiplication, at the second processing element, a second partial sum of the third data matrix based on a second sum of a second product and a second initial value of the third data matrix, wherein the second product is determined based on the first activation bit and the second input data bit; and

streaming, by the system during the matrix multiplication, the first activation bit from the second processing element and along the first dimension, and the second partial sum of the third data matrix from the second processing element and along the second dimension.

5. The computer-implemented method of claim 1 , further comprises:

retaining, by the system during the matrix multiplication, the first initial value of the third data matrix in the first processing element for two clock cycles.

6. The computer-implemented method of claim 1 , wherein the first dimension comprises rows of the systolic array structure and the second dimension comprises columns of the systolic array structure.

7. The computer-implemented method of claim 1 , wherein the first dimension comprises columns of the systolic array structure and the second dimension comprises rows of the systolic array structure.

8. The computer-implemented method of claim 1 , wherein the populating the respective first registers of the processing elements of the systolic array structure occurs in a first time period, and the determining the first partial sum of the third data matrix occurs in a second time period following the first time period.

9. The computer-implemented method of claim 1 , wherein the systolic array structure comprises a rectangular systolic array.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2017
From: CHEN, CHIA-YU; CHOI, JUNGWOOK; GOPALAKRISHNAN, KAILASH; HAN, VICTOR; SRINIVASAN, SRINIVASAN, VIJAYALAKSHMI; ZHANG, JINTAO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044401/0300 →
Continuity (2)
Continuation 15460755 · Mar 16, 2017
Related Publication 20180267938A1 · Sep 20, 2018