IP Library Granted Patent US 11,868,243
Granted Patent B2
US 11,868,243 · App. 17/844,981 · Granted Jan 9, 2024

Topological scheduling

Inventor: Lukasz Lew (Sunnyvale, CA)
Assignee: Google LLC
G06F12/0207G06F7/523G06N3/063G06N20/00G06F2212/2024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,868,243
App. No.
17/844,981
Granted
Jan 9, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing topological scheduling on a machine-learning accelerator having an array of tiles. One of the methods includes performing, at each time step of a plurality of time steps corresponding respectively to columns within each of a plurality of wide columns of the tile array, operations comprising: performing respective multiplications using tiles in a respective tile column for the time step, computing a respective output result for each respective tile column for the time step including computing a sum of results of the multiplications for the tile column, and storing the respective output result for the tile column in a particular output RAM having a location within the same tile column and on a row from which the output result will be read by a subsequent layer of the model.

Claims (37)

1. A device comprising

a tile array comprising a plurality of tiles, wherein the device is configured to perform operations comprising:

receiving a plurality of input activations for a first layer of a model; and

performing, at each time step of a plurality of time steps, operations on tile columns within each tile wide column of a plurality of tile wide columns of the tile array, wherein each tile wide column of the tile array comprises multiple tile columns, the operations comprising:

distributing different feature values of a respective single input activation location along the tile column of a respective tile wide column,

performing respective compute operations using the different feature value for tiles in the respective tile column for the time step,

computing a respective output result for each respective tile column for the time step including computing a sum of results of the compute operations for the tile column, and

storing the respective output result for the tile column in a particular output RAM from which the output result will be read by a subsequent layer of the model.

2. The device of claim 1 , wherein compute operations along a tile column comprise different features of a respective input activation.

3. The device of claim 1 , wherein the operations further comprise aligning input activations along an edge of each tile wide column of the plurality of tile wide columns of the tile array.

4. The device of claim 3 , wherein aligning the input activations comprises aligning the input activations before performing any of the compute operations.

5. The device of claim 1 , wherein the device is a machine-learning accelerator.

6. The device of claim 1 , wherein computing the respective output result is performed by a vector accumulator that is configured to compute an accumulated sum from the results of the compute operations along a single tile column.

7. The device of claim 1 , wherein the device is configured to read each of the plurality of input activations only once.

8. The device of claim 7 , wherein the device is configured to use conveyor hardware to share each input activation with other tiles in a same row of a tile wide column.

9. A method performed by a device comprising

a tile array comprising a plurality of tiles, the method comprising:

receiving a plurality of input activations for a first layer of a model; and

performing, at each time step of a plurality of time steps, operations on tile columns within each tile wide column of a plurality of tile wide columns of the tile array, wherein each tile wide column of the tile array comprises multiple tile columns, the operations comprising:

distributing different feature values of a respective single input activation location along the tile column of a respective tile wide column,

performing respective compute operations using the different feature value for tiles in the respective tile column for the time step,

computing a respective output result for each respective tile column for the time step including computing a sum of results of the compute operations for the tile column, and

storing the respective output result for the tile column in a particular output RAM from which the output result will be read by a subsequent layer of the model.

10. The method of claim 9 , wherein compute operations along a tile column are comprise different features of a respective input activation.

11. The method of claim 9 , further comprising aligning input activations along an edge of each tile wide column of the plurality of tile wide columns of the tile array.

12. The method of claim 11 , wherein aligning the input activations comprises aligning the input activations before performing any of the compute operations.

13. The method of claim 9 , wherein the device is a machine-learning accelerator.

14. The method of claim 9 , wherein computing the respective output result is performed by a vector accumulator that is configured to compute an accumulated sum from the results of the compute operations along a single tile column.

15. The method of claim 9 , comprising reading each of the plurality of input activations only once.

16. The method of claim 15 , further comprising using conveyor hardware to share each input activation with other tiles in a same row of a tile wide column.

17. One or more non-transitory computer storage media encoded with computer program instructions that when executed by a device comprising a tile array comprising a plurality of tiles, causes the device to perform operations comprising:

receiving a plurality of input activations for a first layer of a model; and

performing, at each time step of a plurality of time steps, operations on tile columns within each tile wide column of a plurality of tile wide columns of the tile array, wherein each tile wide column of the tile array comprises multiple tile columns, the operations comprising:

distributing different feature values of a respective single input activation location along the tile column of a respective tile wide column,

performing respective compute operations using the different feature value for tiles in the respective tile column for the time step,

computing a respective output result for each respective tile column for the time step including computing a sum of results of the compute operations for the tile column, and

storing the respective output result for the tile column in a particular output RAM from which the output result will be read by a subsequent layer of the model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2022
From: LEW, LUKASZ
To: GOOGLE LLC
Reel/Frame 060456/0575 →
Continuity (2)
Continuation 16718049 · Dec 17, 2019
Related Publication 20230019367A1 · Jan 19, 2023