Method and system for sequencing artificial intelligence (AI) jobs for execution at AI accelerators
A method for use with an artificial intelligence (AI) sequencer that is at least partially circuitry implemented and is adapted to be coupled to a plurality of AI accelerators via electronic connection, the method comprising: dispatching by the sequencer, using the electronic connection, an initial stage of an AI job having a plurality of stages to at least one of the AI accelerators for execution of the initial stage, wherein the AI job includes multiple AI functions; upon completion of the initial stage, receiving the AI job back at the sequencer; and, thereafter, dispatching by the sequencer, using the electronic connection, a next stage of the AI job to at least one different one of the AI accelerators; wherein the sequencer dispatches each AI function of the AI job at computer speed over the electronic connection to minimize any possible idle time of the AI accelerators.
1 . A method for use with an artificial intelligence (AI) sequencer that is at least partially circuitry implemented and is adapted to be coupled to a plurality of AI accelerators via electronic connection to process an AI job having a plurality of AI functions, the AI sequencer comprising a plurality of schedulers, the method comprising:
receiving, by the AI sequencer, the AI job, the AI job being sent to the AI sequencer for execution by an application;
scheduling, by each of at least two of the plurality of schedulers, respectively, each of at least one of the AI functions of the AI job so that each of the at least one of the AI functions are to be executed by respective ones of the AI accelerators in parallel; and
dispatching by the AI sequencer, using the electronic connection, the each of at least one of the AI functions to respective ones of the AI accelerators for execution in parallel thereby;
wherein each AI accelerator is a dedicated processor configured to process a specific AI function of the AI job; and
wherein the scheduling is performed so that the AI sequencer dispatches each AI function of the AI job over the electronic connection to minimize any possible idle time of the AI accelerators.
2 . The method of claim 1 , wherein the electronic connection is implemented by a network.
3 . The method of claim 1 , wherein results of the AI job are sent by the AI sequencer to the application upon completion of the AI job.
4 . The method of claim 1 , wherein any of the AI accelerators that completes an AI function provides processing results of the AI function to the AI sequencer.
5 . The method of claim 1 , wherein the scheduling is based on a computational graph received by the AI sequencer.
6 . The method of claim 1 , wherein at least one of the AI accelerators belongs to the AI sequencer and at least one other of the AI accelerators belongs to a different AI sequencer and wherein results of the AI job are sent by one of the AI sequencers to the application upon completion of the AI job by all of the AI sequencers.
7 . An artificial intelligence (AI) sequencer adapted to be coupled to a plurality of AI accelerators, the AI sequencer, comprising:
a queue controller comprising logic configured to manage a plurality of queues for maintaining data of a plurality of different AI jobs for each of which a request to employ at least one of the AI accelerators is received, wherein an AI job includes at least one AI functions;
a plurality of schedulers that are each configured to schedule execution of at least a respective one of the AI jobs that is associated with the data maintained by the plurality of queues for the one AI job;
a plurality of job processing units (JPUs), wherein each of the plurality of JPUs is configured to generate an execution sequence for the one of the AI jobs based on a received AI job descriptor for the one of the AI jobs; and
a plurality of dispatchers connected to the plurality of AI accelerators, wherein each of the plurality of dispatchers is configured to dispatch at least an AI function of one of the AI jobs from the plurality of queues to one of the AI accelerators, wherein each AI function is dispatched to one of the AI accelerators in an order determined by the execution sequence created for each respective one of the AI jobs;
wherein the AI sequencer is at least partially circuitry implemented and is couplable to the AI accelerators via electronic connection;
wherein the AI sequencer operates to dispatch each AI function of the AI job to minimize any possible idle time of the AI accelerators; and
wherein, the AI sequencer operates to dispatch at least two of the AI functions of an AI job so that at least at one point in time different ones of the AI accelerators are each executing an AI function of a single AI job in parallel.
8 . The AI sequencer of claim 7 , wherein the electronic connection is implemented by a network.
9 . The AI sequencer of claim 7 , wherein results of a particular AI job are sent by the AI sequencer to an application that requested execution of the particular AI job upon completion of the particular AI job.
10 . The AI sequencer of claim 7 , wherein any of the AI accelerators that completes an AI function provides processing results of the AI function to the AI sequencer.
11 . The AI sequencer of claim 7 , wherein the scheduling is based on a computational graph received by the AI sequencer.
12 . The AI sequencer of claim 7 , wherein the queue controller, the plurality of schedulers, the plurality of JPUs, and the plurality of dispatchers are interconnected by a bus.
13 . The AI sequencer of claim 7 , wherein at least one of the AI accelerators belongs to the AI sequencer and at least one other of the AI accelerators belongs to a different AI sequencer and wherein results of a particular AI job are sent by the AI sequencer to an application that requested execution of the particular AI job upon completion of the particular AI job by all of the AI sequencers.