Hierarchical compiling and execution in a machine learning hardware accelerator
This disclosure describes a system and method for compiling and executing machine learning inferences in an array of multi-core computing devices. Each multi-core computing device can be an application specific integrated circuit (ASIC) or group of ASICS. In many applications, the array of computing devices changes from inference to inference, and can be adjusted based on the requirements of the inference. Additionally, each ASIC can have multiple processing cores, and multiple types of processing cores. Therefore, performing optimizations and scheduling at compile time, can dramatically increase the efficiency of the array in executing the inference. In some implementations, it is possible to select an amount of time or effort to be spent optimizing during compiling, giving the user flexibility in determining whether to spend time during compilation or during execution.
1 . A method for distributing executable jobs in an array of multi-core computing devices, comprising:
receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;
assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:
analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;
assigning the particular job to a core type based on the analysis;
compiling each job of the plurality of jobs into an individually executable file; and
generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,
wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:
a TPU tier, comprising executables to be run on the TPU tile processor;
a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor; and
a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.
2 . The method of claim 1 , further comprising:
executing the individually executable files by:
receiving the execution graph;
assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;
executing, by each multi-core computing device, the assigned jobs;
returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and
combining the returned outputs to generate an execution graph return.
3 . The method of claim 1 , wherein analyzing the particular job is completed using a heuristic analysis.
4 . The method of claim 1 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.
5 . The method of claim 1 , wherein the execution graph comprises
a multi-chip tier, comprising two or more chip-level sub-graphs.
6 . The method of claim 5 , wherein the core of the third type is a host device central processing unit (CPU).
7 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for distributing executable jobs in an array of multi-core computing devices, the operations comprising:
receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;
assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:
analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;
assigning the particular job to a core type based on the analysis;
compiling each job of the plurality of jobs into an individually executable file; and
generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,
wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:
a TPU tier, comprising executables to be run on the TPU tile processor;
a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor; and
a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.
8 . The computer-readable medium of claim 7 , the operations further comprising:
executing the individually executable files by:
receiving the execution graph;
assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;
executing, by each multi-core computing device, the assigned jobs;
returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and
combining the returned outputs to generate an execution graph return.
9 . The computer-readable medium of claim 7 , wherein analyzing the particular job is completed using a heuristic analysis.
10 . The computer-readable medium of claim 7 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.
11 . The computer-readable medium of claim 9 , wherein the execution graph comprises
a multi-chip tier, comprising two or more chip-level sub-graphs.
12 . The computer-readable medium of claim 11 , wherein the core of the third type is a host device central processing unit (CPU).
13 . A system, comprising:
one or more computers; and
a computer-readable storage device coupled to the one or more computers and having instructions stored thereon which, when executed by the one or more computer, cause the one or more computers to perform operations for distributing executable jobs in an array of multi-core computing devices, the operations comprising:
receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device comprising a plurality of different types of processing cores, wherein the plurality of different types of processing cores comprises a core processor and a tensor processing unit (TPU) tile processor;
assigning each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores by:
analyzing the particular job to determine which of the plurality of different types of processing cores is suited to executing the particular job;
assigning the particular job to a core type based on the analysis;
compiling each job of the plurality of jobs into an individually executable file; and
generating an execution graph representing a mapping of the individually executable files to specific ones of the plurality of different types of processing cores, wherein the execution graph identifies dependencies between individually executable files,
wherein the execution graph is hierarchical in nature comprising sub-graphs arranged in at least three tiers comprising:
a TPU tier, comprising executables to be run on the TPU tile processor;
a chip-level tier, comprising one or more sub-graphs of the TPU tier and executables to be run on the core processor;
a host-level tier, comprising a multi-chip tier sub-graph and one or more sub-graphs configured to be executed on a core of a third type.
14 . The system of claim 13 , the operations further comprising:
executing the individually executable files by:
receiving the execution graph;
assigning jobs in the execution graph to a plurality of multi-core computing devices in the array of multi-core computing devices;
executing, by each multi-core computing device, the assigned jobs;
returning, by each multi-core computing device, outputs of the executed jobs to a shared memory; and
combining the returned outputs to generate an execution graph return.
15 . The system of claim 13 , wherein analyzing the particular job is completed using a heuristic analysis.
16 . The system of claim 13 , wherein a depth of analyzation for each particular job is selected based on a user input prior to compile time.