IP Library › Granted Patent US 12,455,762
Granted Patent B2
US 12,455,762 · App. 17/635,772 · Granted Oct 28, 2025

Explicit scheduling of on-chip operations

Inventors: Michial Allen Gunter (Oakland, CA); Charles Henry Leichner, IV (Palo Alto, CA)
Assignee: Google LLC
G06F9/4881G06F9/30036G06F9/321G06F9/3893G06F15/7807G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,455,762
App. No.
17/635,772
Granted
Oct 28, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for obtaining a first schedule, for a first hardware block of an integrated circuit device, where the first schedule identifies a first set of operations to be performed by the first hardware block. Obtaining a second schedule for a second hardware block of the integrated circuit device, where the second schedule identifies a second set of operations to be performed by the second hardware block and where operations of the second schedule are coordinated with operations of the first schedule such that the first schedule triggers the first hardware block to send data to the second block at a first pre-scheduled value of a counter, and the second schedule triggers the second hardware block to accept the data at an input at a second pre-scheduled value of the counter that is after the first pre-scheduled value. Performing, by the first hardware block, the first set of operations according to the first schedule, and performing, by the second hardware block, the second set of operations according to the second schedule.

Claims (36)

1. An integrated circuit device comprising:

a counter;

a first hardware block communicably coupled to the counter and configured to operate according to a first schedule that comprises a first set of operations each of which is scheduled to be executed by the first hardware block at a first respective value of the counter; and

a second hardware block communicably coupled to the counter and to the first hardware block, the second hardware block configured to operate according to a second schedule that comprises a second set of operations each of which is scheduled to be executed by the second hardware block at a second respective value of the counter, and

wherein operations of the second schedule are coordinated with operations of the first schedule such that compute operations in the first schedule are executed concurrently with data exchange operations in the second schedule, and that the first schedule triggers the first hardware block to send data to the second hardware block at a first pre-scheduled value of the counter, and the second schedule triggers the second hardware block to accept the data at an input at a second pre-scheduled value of the counter that is after the first pre-scheduled value.

2. The device of claim 1 , wherein the first set of operations and the second set of operations each comprise a respective portion of a machine learning program.

3. The device of claim 1 , wherein each operation in the first set of operations executes in a predetermined number of clock cycles.

4. The device of claim 1 , wherein operations of the first schedule and the second schedule are coordinated to allow exchange of data between the first hardware block and the second hardware block independent of flow control signals.

5. The device of claim 1 , further comprising a plurality of other hardware blocks, wherein operations of the first schedule are coordinated with respective operation schedules of the other hardware blocks to allow exchange data between the first hardware block and one or more of the other hardware blocks independent of data flow control signals.

6. The device of claim 1 , wherein the first hardware block comprises:

local memory configured to store the first schedule; and

control circuitry coupled to the local memory and configured to execute the first set of operations of the first schedule.

7. The device of claim 6 , wherein the control circuitry is configured to decompress a portion of the first schedule before executing operations included in the portion.

8. The device of claim 1 , wherein the integrated circuit device is an application specific integrated circuit.

9. The device of claim 1 , wherein the first hardware block and the second hardware block are hardware tiles including special purpose circuitry configured to perform neural network operations.

10. The device of claim 9 , wherein the first hardware block comprises:

a computational array of cells; and

local memory coupled to the computational array of cells.

11. The device of claim 1 , wherein the first schedule and the second schedule each comprise a portion of a program executed by the integrated circuit device.

12. An integrated circuit operating method comprising:

obtaining, for a first hardware block of an integrated circuit device, a first schedule that identifies a first set of operations to be performed by the first hardware block;

obtaining, for a second hardware block of the integrated circuit device, a second schedule that identifies a second set of operations to be performed by the second hardware block, wherein operations of the second schedule are coordinated with operations of the first schedule such that compute operations in the first schedule are executed concurrently with data exchange operations in the second schedule, and that the first schedule triggers the first hardware block to send data to the second block at a first pre-scheduled value of a counter, and the second schedule triggers the second hardware block to accept the data at an input at a second pre-scheduled value of the counter that is after the first pre-scheduled value;

performing, by the first hardware block, the first set of operations according to the first schedule; and

performing, by the second hardware block, the second set of operations according to the second schedule.

13. The method of claim 12 , wherein the first schedule and the second schedule each comprise a portion of a program executed by the integrated circuit device.

14. The method of claim 12 , wherein the first set of operations and the second set of operations each comprise a respective portion of a machine learning program.

15. The method of claim 12 , wherein each operation in the first set of operations executes in a predetermined number of clock cycles.

16. The method of claim 12 , wherein operations of the first schedule and the second schedule are coordinated to allow exchange of data between the first hardware block and the second hardware block independent of flow control signals.

17. The method of claim 12 , further comprising decompressing, by the first hardware block, a portion of the first schedule before executing operations included in the portion.

18. The method of claim 12 , wherein the first schedule comprises, for each operation in the first set of operations, a scheduled counter value and data indicating a particular operation to be executed by the first hardware block at the scheduled counter value.

19. The method of claim 12 , wherein performing, by the first hardware block, the first set of operations according to the first schedule comprises:

receiving, from a counter, a first counter value that equals a first scheduled counter value of a first operation in the first set of operations;

in response to receiving the first counter value, causing a first set of one or more computational units of the first hardware block to execute the first operation;

receiving, from the counter, a second counter value that equals a second scheduled counter value of a second operation in the first set of operations; and

in response to receiving the second counter value, causing a second set of one or more computational units of the first hardware block to execute the second operation.

20. The method of claim 12 , wherein the first hardware block and the second hardware block are hardware tiles including special purpose circuitry configured to perform neural network operations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: GUNTER, MICHIAL ALLEN; LEICHNER IV, CHARLES HENRY
To: GOOGLE LLC
Reel/Frame 059197/0390 →
Continuity (2)
Provisional Application 62887724 · Aug 16, 2019
Related Publication 20220326988A1 · Oct 13, 2022
References Cited (24)
US 20150186313A1 · Sodhi et al. · 2015 [cited by applicant]
US 20190114548A1 · Wu et al. · 2019 [cited by applicant]
US 20190121388A1 · Knowles et al. · 2019 [cited by applicant]
US 20190121784A1 · Wilkinson et al. · 2019 [cited by applicant]
US 20190340023A1 · Brewer · 2019 [cited by examiner]
CN 108268386 · 2018 [cited by applicant]
CN 108573304 · 2018 [cited by applicant]
CN 108694109 · 2018 [cited by applicant]
JP 2016511853 · 2016 [cited by applicant]
JP 2018101917 · 2018 [cited by applicant]
JP 2019079522 · 2019 [cited by applicant]
JP 2019079526 · 2019 [cited by applicant]
TW 201826122 · 2018 [cited by applicant]
TW 201901534 · 2019 [cited by applicant]
WO WO2014099539 · 2014 [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2022-509681, dated Apr. 17, 2023, 9 pages (with English translation). [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2022-509681, dated Jul. 31, 2023, 5 pages (with English translation). [cited by applicant]
International Preliminary Report on Patentability International Appln. No. PCT/US2020/046392, dated Feb. 17, 2022, 12 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2020/046392, dated Nov. 5, 2020, 15 pages. [cited by applicant]
Liang et al., “An Architecture and Compiler for Scalable On-Chip Communication,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Jul. 2004, 12(7):711-726. [cited by applicant]
Office Action in Taiwan Appln. No. 109127774, dated Aug. 17, 2021, 21 pages (with English translation). [cited by applicant]
Notice of Allowance in Chinese Appln. No. 202080057906.0, mailed on Jan. 6, 2024, 7 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7007685, mailed on Dec. 8, 2023, 20 pages (with English translation). [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2022-7007685, mailed on Oct. 30, 2024, 7 pages (with English translation). [cited by applicant]