IP Library › Granted Patent US 12,223,211
Granted Patent B2
US 12,223,211 · App. 18/468,630 · Granted Feb 11, 2025

Enhanced input of machine-learning accelerator activations

Inventors: Lukasz Lew (Sunnyvale, CA); Wren Romano (Mountain View, CA)
Assignee: Google LLC
G06F3/0679G06F3/0604G06F3/0655G06F9/30036G06F9/50G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,211
App. No.
18/468,630
Granted
Feb 11, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for scheduling operations on a machine-learning accelerator having multiple tiles. The apparatus includes a processor having a plurality of tiles and scheduling circuitry that is configured to select a respective input activation for each tile of the plurality of tiles from either an activation line for the tile or a delay register for the activation line.

Claims (16)

1. A method performed using a processor comprising a plurality of activation lines and a plurality of delay registers, wherein each delay register is paired with a respective activation line, the method comprising:

on a first global tick, populating a first delay register with a first activation value from a first activation line of the processor, and

routing a second activation value to the first activation line;

on a first local tick of the first global tick, providing at least the first activation value from the delay register to a compute tile of the processor; and

on each subsequent local tick of the first global tick, providing one or more additional activation values from the plurality of activation lines after providing at least the first activation value from the delay register.

2. The method of claim 1 , wherein on each local tick of the first global tick, performing, by the processor, a mathematical operation between an input activation value and a respective weight to generate a partial sum value.

3. The method of claim 2 , wherein the mathematical operation is part of a convolution operation, and wherein the respective weights are from a convolution kernel.

4. The method of claim 2 , comprising computing, by the processor, an accumulated sum at least in part from the partial sum value and an additional partial sum value.

5. The method of claim 1 , further comprising, on a second global tick, populating one or more delay registers with one or more third activation values from the plurality of activation lines, and reading, from a memory device, one or more fourth activation values onto the plurality of activation lines.

6. The method of claim 5 , wherein the memory device comprises static random-access memory.

7. The method of claim 1 , comprising determining each activation value based on a stride length.

8. The method of claim 7 , wherein the stride length is based on a size of a convolution kernel.

9. The method of claim 1 , wherein providing one or more additional activation values from the plurality of activation lines comprises:

determining whether there are more partial sums for the first global tick; and

computing a subsequent partial sum with fewer delay registers when there are more partial sums for the first global tick, or determining whether there are more global ticks when there are no more partial sums for the first global tick.

10. The method of claim 9 , comprising overwriting the plurality of delay registers when there are more global ticks, or ending a process when there are no more global ticks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: LEW, LUKASZ; ROMANO, WREN
To: GOOGLE LLC
Reel/Frame 064927/0888 →
Continuity (4)
Continuation 17738403 · May 6, 2022
Continuation 16718055 · Dec 17, 2019
Provisional Application 62935038 · Nov 13, 2019
Related Publication 20240192897A1 · Jun 13, 2024
References Cited (11)
US 9805303B2 · Ross · 2017 [cited by applicant]
US 10790828B1 · Gunter · 2020 [cited by examiner]
US 20180267898A1 · Henry et al. · 2018 [cited by applicant]
US 20190079761A1 · Diamon et al. · 2019 [cited by applicant]
US 20190121668A1 · Knowles · 2019 [cited by examiner]
US 20220180158A1 · Whatmough · 2022 [cited by examiner]
US 20220197564A1 · Visconti · 2022 [cited by examiner]
CN 109697111 · 2019 [cited by applicant]
Misra et al., “HOG and Spatial Convolution on SIMD Architecture,” Robotics Institute, 2015, 7 pages. [cited by applicant]
Office Action in European Appln. No. 20193932.9, Oct. 6, 2022, 9 pages. [cited by applicant]
Office Action in Taiwan Appln. No. 111117325, dated Sep. 28, 2023, 20 pages (with English translation). [cited by applicant]