IP Library Granted Patent US 11,762,602
Granted Patent B2
US 11,762,602 · App. 17/738,403 · Granted Sep 19, 2023

Enhanced input of machine-learning accelerator activations

Inventors: Lukasz Lew (Sunnyvale, CA); Wren Romano (Mountain View, CA)
Assignee: Google LLC
G06F3/0679G06F3/0604G06F3/0655G06F9/50G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,762,602
App. No.
17/738,403
Granted
Sep 19, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for scheduling operations on a machine-learning accelerator having multiple tiles. The apparatus includes a processor having a plurality of tiles and scheduling circuitry that is configured to select a respective input activation for each tile of the plurality of tiles from either an activation line for the tile or a delay register for the activation line.

Claims (22)

1. A method performed by a processor comprising:

a plurality of tiles;

a plurality of activation lines that couple a memory device to a respective tile, wherein the memory device stores an array of input data;

a plurality of delay registers that couple to a respective tile, wherein each delay register is paired with a corresponding activation line;

the method comprising:

on a first global tick, populating the plurality of delay registers with a plurality of first input activation values from the plurality of activation lines, respectively, and reading, from the memory device, a plurality of second input activation values onto the plurality of activation lines, respectively,

on a first local tick of the first global tick, providing input activation values only from the plurality of delay registers to the plurality of tiles, respectively, and

on each subsequent local tick of the first global tick, providing input activation values from an increasing number of activation lines and a decreasing number of delay registers.

2. The method of claim 1 , wherein on each local tick of the first global tick, performing, by each tile, a mathematical operation between an input activation value received at the tile and a respective weight to generate a partial sum value.

3. The method of claim 1 , wherein the processor is a machine-learning accelerator.

4. The method of claim 1 , wherein the mathematical operation is part of a convolution operation, and wherein the respective weights are from a convolution kernel.

5. The method of claim 2 , further comprising computing, by a vector accumulator, an accumulated sum from partial sum values computed by each tile of the plurality of tiles.

6. The method of claim 1 , further comprising, on a second global tick, populating the plurality of delay registers with the plurality of second input activation values from the plurality of activation lines, respectively, and reading, from the memory device, a plurality of third input activation values onto the plurality of activation lines, respectively.

7. The method of claim 1 , wherein the plurality of tiles are arranged in multiple row groups, and each input activation value is read from the memory device only once per row group.

8. The method of claim 1 , comprising determining the input activation values from memory based on a stride length.

9. The method of claim 8 , wherein the stride length is based on a size of a convolution kernel.

10. The method of claim 1 , wherein the memory comprises static random-access memory.

11. The method of claim 1 , wherein providing input activation values from an increasing number of activation lines and a decreasing number of delay registers comprises:

determining whether there are more partial sums for the first global tick; and

computing a subsequent partial sum with fewer delay registers when there are more partial sums for the first global tick, or

determining whether there are more global ticks when there are no more partial sums for the first global tick.

12. The method of claim 11 , comprising overwriting the plurality of delay registers when there are more global ticks, or ending the process when there are no more global ticks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2022
From: LEW, LUKASZ; ROMANO, WREN
To: GOOGLE LLC
Reel/Frame 059841/0934 →
Continuity (3)
Continuation 16718055 · Dec 17, 2019
Provisional Application 62935038 · Nov 13, 2019
Related Publication 20220334776A1 · Oct 20, 2022