IP Library › Granted Patent US 12,632,237
Granted Patent B2
US 12,632,237 · App. 18/561,590 · Granted May 19, 2026

Hierarchical compiling and execution in a machine learning hardware accelerator

Inventors: John Navil Joseph (Kirkland, WA); Jack Liu (Saratoga, CA); Dong Hyuk Woo (San Jose, CA); Jing Pu (Santa Clara, CA)
Assignee: Google LLC
G06F8/451G06F9/451
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,237
App. No.
18/561,590
Filed
Nov 16, 2023
Granted
May 19, 2026
Kind
B2
Art Unit
2199
USPC
717/149
Abstract

This disclosure describes a system and method for compiling and executing machine learning inferences in an array of multi-core computing devices. Each multi-core computing device can be an application specific integrated circuit (ASIC) or group of ASICS. In many applications, the array of computing devices changes from inference to inference, and can be adjusted based on the requirements of the inference. Additionally, each ASIC can have multiple processing cores, and multiple types of processing cores. Therefore, performing optimizations and scheduling at compile time, can dramatically increase the efficiency of the array in executing the inference. In some implementations, it is possible to select an amount of time or effort to be spent optimizing during compiling, giving the user flexibility in determining whether to spend time during compilation or during execution.

Claims (68)

1 . A method for distributing executable jobs in an array of multi-core computing devices, comprising:

receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;

assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:

analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;

assigning the particular job to a core type based on the analysis;

compiling each job of the plurality of jobs into an individually executable file; and

generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,

wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:

a TPU tier, comprising executables to be run on the TPU tile processor;

a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor; and

a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.

2 . The method of claim 1 , further comprising:

executing the individually executable files by:

receiving the execution graph;

assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;

executing, by each multi-core computing device, the assigned jobs;

returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and

combining the returned outputs to generate an execution graph return.

3 . The method of claim 1 , wherein analyzing the particular job is completed using a heuristic analysis.

4 . The method of claim 1 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.

5 . The method of claim 1 , wherein the execution graph comprises

a multi-chip tier, comprising two or more chip-level sub-graphs.

6 . The method of claim 5 , wherein the core of the third type is a host device central processing unit (CPU).

7 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for distributing executable jobs in an array of multi-core computing devices, the operations comprising:

receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;

assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:

analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;

assigning the particular job to a core type based on the analysis;

compiling each job of the plurality of jobs into an individually executable file; and

generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,

wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:

a TPU tier, comprising executables to be run on the TPU tile processor;

a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor; and

a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.

8 . The computer-readable medium of claim 7 , the operations further comprising:

executing the individually executable files by:

receiving the execution graph;

assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;

executing, by each multi-core computing device, the assigned jobs;

returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and

combining the returned outputs to generate an execution graph return.

9 . The computer-readable medium of claim 7 , wherein analyzing the particular job is completed using a heuristic analysis.

10 . The computer-readable medium of claim 7 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.

11 . The computer-readable medium of claim 9 , wherein the execution graph comprises

a multi-chip tier, comprising two or more chip-level sub-graphs.

12 . The computer-readable medium of claim 11 , wherein the core of the third type is a host device central processing unit (CPU).

13 . A system, comprising:

one or more computers; and

a computer-readable storage device coupled to the one or more computers and having instructions stored thereon which, when executed by the one or more computer, cause the one or more computers to perform operations for distributing executable jobs in an array of multi-core computing devices, the operations comprising:

receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;

assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:

analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;

assigning the particular job to a core type based on the analysis;

compiling each job of the plurality of jobs into an individually executable file; and

generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,

wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:

a TPU tier, comprising executables to be run on the TPU tile processor;

a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor;

a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.

14 . The system of claim 13 , the operations further comprising:

executing the individually executable files by:

receiving the execution graph;

assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;

executing, by each multi-core computing device, the assigned jobs;

returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and

combining the returned outputs to generate an execution graph return.

15 . The system of claim 13 , wherein analyzing the particular job is completed using a heuristic analysis.

16 . The system of claim 13 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2023
From: JOSEPH, JOHN NAVIL; LIU, JACK; WOO, DONG HYUK; PU, JING
To: GOOGLE LLC
Reel/Frame 065621/0001 →
Continuity (1)
Related Publication 20240256239A1 · Aug 1, 2024
References Cited (23)
US 10841810B2 · O'Shea · 2020 [cited by examiner]
US 11003429B1 · Zejda et al. · 2021 [cited by applicant]
US 11392845B2 · Singh · 2022 [cited by examiner]
US 12189629B2 · Sen · 2025 [cited by examiner]
US 20170177415A1 · Dhanraj et al. · 2017 [cited by applicant]
US 20190391796A1 · Brady et al. · 2019 [cited by applicant]
US 20200342286A1 · Zhang · 2020 [cited by applicant]
US 20220043688A1 · Lai · 2022 [cited by examiner]
US 20230259774A1 · Wang · 2023 [cited by examiner]
US 20240256333A1 · Weber · 2024 [cited by examiner]
WO WO2020052241A1 · 2020 [cited by applicant]
Ahmed et al., “Adaptive resource management for simultaneous multitasking in mixed-grained reconfigurable multi-core processors,” Paper, Presented at the 7th IEEE/ACM/IFIP international conference on Hardware/software c… [cited by applicant]
Ambrosi et al., “Hardware-software co-design for an analog-digital accelerator for machine learning,” Paper, Presented at the 2018 International Conference on Rebooting Computing, Mclean, VA, Nov. 7-9, 2018; Proceedings… [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2021/036418, mailed Dec. 21, 2023, 12 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2021/036418, mailed on Mar. 28, 2022, 20 pages. [cited by applicant]
Invitation to Pay Additional Fees in International Appln. No. PCT/US2021/036418, mailed on Feb. 7, 2022, 12 pages. [cited by applicant]
Jozwiak et al., “Modern development methods and tools for embedded reconfigurable systems: a survey,” Integration, The VLSI Journal, Jan. 1, 2010, 43(1):1-33. [cited by applicant]
Kourtis et al., “Compiling neural networks for a computational memory accelerator,” Paper, Presented at the 10th Workshop on Systems for Post-Moore Architectures, Heraklion, Greece, Apr. 27, 2020, pp. 1-8. [cited by applicant]
Nollet et al., “Run-time management of a MPSoC containing FPGA fabric tiles,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Jan. 1, 2008, 16(1):24-33. [cited by applicant]
Singh et al., “Mapping on multi/many-core systems: survey of current and emerging trends,” Paper, Presented at the 50th Annual Design Automation Conference, Austin, TX, May 29-Jun. 7, 2013; Proceedings of the 50th Annua… [cited by applicant]
Office Action in Japanese Appln. No. 2023-571345, mailed on Apr. 22, 2025, 6 pages (with English translation). [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2023-571345, mailed on Jun. 17, 2025, 5 pages (with English translation). [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2023-7038781, mailed on Jan. 27, 2026, 4 pages (with English translation). [cited by applicant]