IP Library › Granted Patent US 11,907,330
Granted Patent B2
US 11,907,330 · App. 18/111,468 · Granted Feb 20, 2024

Low latency matrix multiply unit

Inventors: Andrew Everett Phelps (Middleton, WI); Norman Paul Jouppi (Palo Alto, CA)
Assignee: Google LLC
G06F17/16G06F5/015G06F9/30101G06F15/8046G06F9/30032G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,907,330
App. No.
18/111,468
Granted
Feb 20, 2024
Kind
B2
Abstract

Methods, systems, and apparatus for a matrix multiply unit implemented as a systolic array of cells are disclosed. Each cell of the matrix multiply includes: a weight matrix register configured to receive a weight input from either a transposed or a non-transposed weight shift register; a transposed weight shift register configured to receive a weight input from a horizontal direction to be stored in the weight matrix register; a non-transposed weight shift register configured to receive a weight input from a vertical direction to be stored in the weight matrix register; and a multiply unit that is coupled to the weight matrix register and configured to multiply the weight input of the weight matrix register with a vector data input in order to obtain a multiplication result.

Claims (55)

1. A processor comprising:

one or more cores, each core having an array of cells, and

one or more storage devices on which are stored instructions that are operable, when executed by a core of the one or more cores, to cause the core to perform operations comprising:

receiving data representing instructions for performing a portion of operations;

generating instructions that, for each cell of an array of cells included in the core and once executed by the cell, causes the cell to perform operations comprising:

receiving, by a weight matrix register of the cell, a weight input from a plurality of weight storing registers, wherein each of the plurality of weight storing registers is configured to receive a respective weight input from a respective direction of the array, wherein the respective directions comprise a first direction of the array and a second direction of the array different from the first direction; and

obtaining, by a multiply unit that is coupled to the weight matrix register, a respective multiplication result based on the received weight input and another input; and

generating a result for performing the portion of operations by combining the respective multiplication results obtained from the array of cells.

2. The processor of claim 1 , wherein the portion of operations is associated with a network layer of a neural network, the respective weight input is associated with the network layer, and the other input includes an activation input of the network layer.

3. The processor of claim 2 , wherein the operations performed by the cell further comprises:

receiving, by the multiply unit, another weight input from the weight matrix register; and

multiplying, by the multiply unit, the other received weight input with another activation input associated with the network layer to generate another multiplication result.

4. The processor of claim 1 , wherein the array of cells is a systolic array of cells, and the array has a two-dimensional format, wherein the first direction of the array is the first direction in the two-dimensional format, and the second direction of the array is the second direction in the two-dimensional format.

5. The processor of claim 1 , wherein receiving, by the weight matrix register of the cell, the weight input from the plurality of weight storing registers further comprises:

selecting, by a multiplexer of the cell, the weight input from at least two respective weight inputs each stored in a corresponding weight storing register of the plurality of weight storing registers; and

receiving, by the weight matrix register, the selected weight input from the multiplexer.

6. The processor of claim 1 , wherein the plurality of weight storing registers comprise a transposed weight shift register and a non-transposed weight shift register physically separate from the transposed weight shift register.

7. The processor of claim 1 , wherein the plurality of weight storing registers comprise:

a first weight storing register configured to receive a first weight input over a first wired path from a first cell of the array of cells that is along the first direction; and

a second weight storing register configured to receive a second weight input over a second wired path from a second cell of the array of cells that is along the second direction;

wherein the weight input is one of the first weight input and the second weight input.

8. The processor of claim 1 , wherein the core further includes one or more of a scalar memory, a vector memory, a scalar processing unit, one or more vector registers, a transposed unit, or a reduction and permutation unit.

9. A method performed by a core of one or more cores included in a processor, the method comprising:

receiving data representing instructions for performing a portion of operations;

generating instructions that, for each cell of an array of cells included in the core and once executed by the cell, causes the cell to perform operations comprising:

receiving, by a weight matrix register of the cell, a weight input from a plurality of weight storing registers, wherein each of the plurality of weight storing registers is configured to receive a respective weight input from a respective direction of the array, wherein the respective directions comprise a first direction of the array and a second direction of the array different from the first direction; and

obtaining, by a multiply unit that is coupled to the weight matrix register, a respective multiplication result based on the received weight input and another input; and

generating a result for performing the portion of operations by combining the respective multiplication results obtained from the array of cells.

10. The method of claim 9 , wherein the portion of operations is associated with a network layer of a neural network, the respective weight input is associated with the network layer, and the other input includes an activation input of the network layer.

11. The method of claim 10 , wherein the operations performed by the cell further comprises:

receiving, by the multiply unit, another weight input from the weight matrix register; and

multiplying, by the multiply unit, the other received weight input with another activation input associated with the network layer to generate another multiplication result.

12. The method of claim 10 , wherein the array of cells is a systolic array of cells, and the array has a two-dimensional format, wherein the first direction of the array is the first direction in the two-dimensional format, and the second direction of the array is the second direction in the two-dimensional format.

13. The method of claim 10 , wherein receiving, by the weight matrix register of the cell, the weight input from the plurality of weight storing registers further comprises:

selecting, by a multiplexer of the cell, the weight input from at least two respective weight inputs each stored in a corresponding weight storing register of the plurality of weight storing registers; and

receiving, by the weight matrix register, the selected weight input from the multiplexer.

14. The method of claim 10 , wherein the plurality of weight storing registers comprise a transposed weight shift register and a non-transposed weight shift register physically separate from the transposed weight shift register.

15. The method of claim 10 , wherein the plurality of weight storing registers comprise:

a first weight storing register configured to receive a first weight input over a first wired path from a first cell of the array of cells that is along the first direction; and

a second weight storing register configured to receive a second weight input over a second wired path from a second cell of the array of cells that is along the second direction;

wherein the weight input is one of the first weight input and the second weight input.

16. The method of claim 10 , wherein the core further includes one or more of a scalar memory, a vector memory, a scalar processing unit, one or more vector registers, a transposed unit, or a reduction and permutation unit.

17. A non-transitory computer program product storing instructions that, when executed by a core of one or more cored included in a programmable processor, cause the core to perform operations comprising:

receiving data representing instructions for performing a portion of operations;

generating instructions that, for each cell of an array of cells included in the core and once executed by the cell, causes the cell to perform operations comprising:

receiving, by a weight matrix register of the cell, a weight input from a plurality of weight storing registers, wherein each of the plurality of weight storing registers is configured to receive a respective weight input from a respective direction of the array, wherein the respective directions comprise a first direction of the array and a second direction of the array different from the first direction; and

obtaining, by a multiply unit that is coupled to the weight matrix register, a respective multiplication result based on the received weight input and another input; and

generating a result for performing the portion of operations by combining the respective multiplication results obtained from the array of cells.

18. The non-transitory computer program product of claim 17 , wherein the portion of operations is associated with a network layer of a neural network, the respective weight input is associated with the network layer, and the other input includes an activation input of the network layer.

19. The non-transitory computer program product of claim 18 , wherein the operations performed by the cell further comprises:

receiving, by the multiply unit, another weight input from the weight matrix register; and

multiplying, by the multiply unit, the other received weight input with another activation input associated with the network layer to generate another multiplication result.

20. The non-transitory computer program product of claim 17 , wherein receiving, by the weight matrix register of the cell, the weight input from the plurality of weight storing registers further comprises:

selecting, by a multiplexer of the cell, the weight input from at least two respective weight inputs each stored in a corresponding weight storing register of the plurality of weight storing registers; and

receiving, by the weight matrix register, the selected weight input from the multiplexer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: PHELPS, ANDREW EVERETT; JOUPPI, NORMAN PAUL
To: GOOGLE INC.
Reel/Frame 065011/0434 →
CERTIFICATE OF CONVERSION Recorded Sep 25, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 065022/0485 →
Continuity (6)
Continuation 17210293 · Mar 23, 2021
Continuation 16915286 · Jun 29, 2020
Continuation 16529662 · Aug 1, 2019
Continuation 15983037 · May 17, 2018
Provisional Application 62507766 · May 17, 2017
Related Publication 20230267172A1 · Aug 24, 2023