IP Library › Granted Patent US 11,853,865
Granted Patent B2
US 11,853,865 · App. 17/134,936 · Granted Dec 26, 2023

Prefetching weights for use in a neural network processor

Inventor: Jonathan Ross (Mountain View, CA)
Assignee: Google LLC
G06N3/063G06F15/8046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,853,865
App. No.
17/134,936
Granted
Dec 26, 2023
Kind
B2
Abstract

A circuit for performing neural network computations for a neural network, the circuit comprising: a systolic array comprising a plurality of cells; a weight fetcher unit configured to, for each of the plurality of neural network layers: send, for the neural network layer, a plurality of weight inputs to cells along a first dimension of the systolic array; and a plurality of weight sequencer units, each weight sequencer unit coupled to a distinct cell along the first dimension of the systolic array, the plurality of weight sequencer units configured to, for each of the plurality of neural network layers: shift, for the neural network layer, the plurality of weight inputs to cells along the second dimension of the systolic array over a plurality of clock cycles and where each cell is configured to compute a product of an activation input and a respective weight input using multiplication circuitry.

Claims (52)

1. A system for performing neural network computations for a neural network having a plurality of neural network layers, the system comprising:

a matrix computation unit comprising an array of cells wherein each cell includes respective circuitry configured to:

obtain a respective weight input for a neural network layer of the plurality of neural network layers;

receive a control signal from a sequencer; and

determine, based on the control signal, whether to reuse the respective weight input for another cell at a subsequent clock cycle.

2. The system of claim 1 , wherein each cell is further configured to:

in response to a determination to reuse the respective weight input for the other cell at the subsequent clock cycle, shift the respective weight input to the other cell.

3. The system of claim 1 , wherein each cell is further configured to:

obtain a respective activation input for the neural network layer; and

determine, based on the control signal, Whether to reuse the respective activation input.

4. The system of claim 3 , wherein:

the array of cells comprises a first dimension and a second dimension; and

for at least one cell along the first dimension, obtaining the respective activation input comprises obtaining the respective activation input from a respective value loader.

5. The system of claim 3 , wherein for at least one cell in the array of cells:

obtaining the respective weight input comprises obtaining the respective weight input shifted from a respective first cell of the array of cell; and

obtaining the respective activation input comprises obtaining the respective activation input shifted from a respective second cell of the array of cells that is different from the respective first cell.

6. The system of claim 3 , further comprising:

a first memory configured to provide activation inputs for the plurality of neural network layers; and

a second memory configured to provide weight inputs for the plurality of neural network layers.

7. The system of claim 6 , further comprising vector computation circuitry configured to:

determine a vector based on one or more accumulated values received from the matrix computation unit; and

provide the vector to the first memory.

8. The system of claim 6 , further comprising:

sequencer circuitry configured to provide one or more control signals to the first memory, the second memory, the vector computation circuitry, or the matrix computation unit to control a dataflow of the system.

9. The system of claim 1 , wherein:

the array of cells comprises a first dimension and a second dimension; and

for at least one cell along the second dimension, obtaining the respective weight input comprises obtaining the respective weight input from a weight fetcher interface.

10. The system of claim 1 , wherein determining, based on the control signal, whether to reuse the respective weight input comprises determining that the control signal is equal to a predetermined value.

11. The system of claim 10 , wherein the respective circuitry of each cell comprises:

a respective weight control register configured to store the control signal; and

a respective weight register configured to load the respective weight input.

12. The system of claim 1 , wherein the matrix computation unit further comprises a plurality of sequencers, each sequencer configured to provide a respective control signal to a corresponding cell of the array of cells.

13. The system of claim 12 , wherein each sequencer includes decrement circuitry configured to decrement, at each clock cycle, a value of the respective control signal provided to the corresponding cell.

14. The system of claim 12 , wherein one or more sequencers of the plurality of sequencers are configured to provide the respective control signal to a subsequent sequencer of the plurality of sequencers.

15. The system of claim 1 , wherein the array of cells form a systolic array.

16. A method for performing neural network computations for a neural network having a plurality of neural network layers, the method comprising:

obtaining, by a matrix computation unit comprising an array of cells, a respective weight input for a neural network layer of the plurality of neural network layers;

obtaining, by each cell in the array of cells, a respective weight input for the neural network layer;

receiving, by each cell, a control signal from a sequencer; and

determining, by each cell and based on the control signal, whether to reuse the respective weight input for another cell at a subsequent clock cycle.

17. The method of claim 16 , further comprising:

in response to determining to reuse the respective weight input, shifting the respective weight input to the other cell.

18. The method of claim 16 , further comprising:

obtaining, by each cell, a respective activation input for the neural network layer; and

determining, by each cell and based on the control signal, whether to reuse the respective activation input.

19. A matrix computation unit for performing neural network computations for a neural network having a plurality of neural network layers, the matrix computation unit comprising:

an array of cells, wherein each cell includes respective circuitry configured to:

obtain a respective weight input for a neural network layer of the plurality of neural network layers;

receive a control signal from a sequencer; and

determine, based on the control signal, whether to reuse the respective weight input for another cell at a subsequent clock cycle.

20. The matrix computation unit of claim 19 , wherein each cell is further configured to:

in response to a determination to reuse the respective weight input for the other cell at the subsequent clock cycle, shift the respective weight input to the other cell.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2020
From: ROSS, JONATHAN
To: GOOGLE INC.
Reel/Frame 054783/0808 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 054882/0924 →
Continuity (5)
Continuation 16826466 · Mar 23, 2020
Continuation 16053305 · Aug 2, 2018
Continuation 14844670 · Sep 3, 2015
Provisional Application 62164981 · May 21, 2015
Related Publication 20210192328A1 · Jun 24, 2021