IP Library › Granted Patent US 10,963,176
Granted Patent B2
US 10,963,176 · App. 15/721,821 · Granted Mar 30, 2021

Technologies for offloading acceleration task scheduling operations to accelerator sleds

Inventors: Susanne M. Balle (Hudson, NH); Francesc Guim Bernat (Barcelona, ES); Slawomir Putyrski (Gdynia, PL); Joe Grecco (Saddle Brook, NJ); Henry Mitchel (Wayne, NJ); Rahul Khanna (Portland, OR); Evan Custodio (Seekonk, MA)
Assignee: Intel Corporation
G06F3/0641G06F3/0604G06F3/065G06F3/067G06F3/0608G06F3/0611G06F3/0613G06F3/0617G06F3/0647G06F3/0653G06F7/06G06F8/65G06F8/654G06F8/656G06F8/658G06F9/3851G06F9/3891G06F9/4401G06F9/4881G06F9/505G06F9/5005G06F9/5038G06F9/544G06F11/0709G06F11/079G06F11/0751G06F11/3006G06F11/3034G06F11/3055G06F11/3079G06F11/3409G06F12/0284G06F12/0692G06F13/1652G06F16/1744G06F21/57G06F21/6218G06F21/73G06F21/76G06T1/20G06T1/60G06T9/005H01R13/453H01R13/4536H01R13/4538H01R13/631H03K19/1731H03M7/3084H03M7/40H03M7/42H03M7/60H03M7/6011H03M7/6017H03M7/6029H04L9/0822H04L12/2881H04L12/4633H04L41/044H04L41/0816H04L41/0853H04L41/12H04L43/04H04L43/06H04L43/08H04L43/0894H04L47/20H04L47/2441H04L49/104H04L61/2007H04L67/10H04L67/1014H04L67/327H04L67/36H05K7/1452H05K7/1487G06F11/1453G06F12/023G06F15/80G06F2212/401G06F2212/402G06F2221/2107H04L41/046H04L41/0896H04L41/142H04L47/78H04L63/1425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,963,176
App. No.
15/721,821
Granted
Mar 30, 2021
Kind
B2
Abstract

Technologies for offloading acceleration task scheduling operations to accelerator sleds include a compute device to receive a request from a compute sled to accelerate the execution of a job, which includes a set of tasks. The compute device is also to analyze the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks. Additionally the compute device is to send an availability request, including the metadata, to one or more micro-orchestrators of one or more accelerator sleds communicatively coupled to the compute device. The compute device is further to receive availability data from the one or more micro-orchestrators, indicative of which of the tasks the micro-orchestrator has accepted for acceleration on the associated accelerator sled. Additionally, the compute device is to assign the tasks to the one or more micro-orchestrators as a function of the availability data.

Claims (55)

1. A compute device comprising:

an I/O subsystem and

circuitry to:

receive a request from a compute sled to accelerate execution of a job, wherein the job includes a set of tasks;

analyze the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks;

send an availability request to a micro-orchestrator of an accelerator sled communicatively coupled to the compute device, wherein the availability request includes the metadata;

receive availability data from the micro-orchestrator, wherein the availability data is indicative of which of the tasks the micro-orchestrator has accepted for acceleration on an associated accelerator sled; and

assign the tasks to the micro-orchestrator as a function of the availability data.

2. The compute device of claim 1 , wherein to receive a request to accelerate a job comprises to receive a request that includes code indicative of operations to be performed within the job; and

wherein to analyze the request comprises to analyze the code to identify operations to be grouped into tasks.

3. The compute device of claim 1 , wherein to analyze the request comprises to determine a type of acceleration for each task.

4. The compute device of claim 1 , wherein to analyze the request comprises to determine a data dependence between the tasks.

5. The compute device of claim 4 , wherein to determine the data dependence between the tasks comprises determine a subdivision of the tasks to operate on different portions of a data set concurrently.

6. The compute device of claim 1 , wherein to receive the request to accelerate a job comprises to receive a request that identifies a workload phase associated with the job; and

wherein the circuitry is further to associate the tasks within the job with a workload phase identifier.

7. The compute device of claim 1 , wherein to send an availability request to a micro-orchestrator comprises to send the availability request to multiple micro-orchestrators.

8. The compute device of claim 1 , wherein to receive availability data from the micro-orchestrator comprises to receive an indication of an estimated time to complete the tasks accepted by a micro-orchestrator.

9. The compute device of claim 1 , wherein to receive availability data from the micro-orchestrator comprises to receive an indication of whether an accelerator sled associated with the micro-orchestrator can access a shared memory with another accelerator sled for parallel execution of one or more of the tasks.

10. The compute device of claim 1 , wherein the accelerator sled is one of a plurality of accelerator sleds and wherein to assign the tasks to the micro-orchestrator comprises to assign the tasks as a function of a best fit of each associated accelerated sled to the tasks.

11. The compute device of claim 10 , wherein to assign the tasks as a function of a best fit of each associated accelerated sled comprises to consolidate tasks on the associated accelerator sleds to reduce network congestion.

12. The compute device of claim 10 , wherein to assign the tasks as a function of a best fit of each associated accelerated sled comprises to assign the tasks as a function of estimated time completion of each task.

13. One or more non-transitory machine-readable storage media comprising a plurality of instructions stored thereon that, in response to being executed, cause a compute device to:

receive a request from a compute sled to accelerate execution of a job, wherein the job includes a set of tasks;

analyze the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks;

send an availability request to a micro-orchestrator of an accelerator sled communicatively coupled to the compute device, wherein the availability request includes the metadata;

receive availability data from the micro-orchestrator, wherein the availability data is indicative of which of the tasks the micro-orchestrator has accepted for acceleration on an associated accelerator sled; and

assign the tasks to the micro-orchestrator as a function of the availability data.

14. The one or more non-transitory machine-readable storage media of claim 13 , wherein to receive a request to accelerate a job comprises to receive a request that includes code indicative operations to be performed within the job; and

wherein to analyze the request comprises to analyze the code to identify operations to be grouped into tasks.

15. The one or more non-transitory machine-readable storage media of claim 13 , wherein to analyze the request comprises to determine a type of acceleration for each task.

16. The one or more non-transitory machine-readable storage media of claim 13 , wherein to analyze the request comprises to determine a data dependence between the tasks.

17. The one or more non-transitory machine-readable storage media of claim 16 , wherein to determine the data dependence between the tasks comprises determine a subdivision of the tasks to operate on different portions of a data set concurrently.

18. The one or more non-transitory machine-readable storage media of claim 13 , wherein to receive the request to accelerate a job comprises to receive a request that identifies a workload phase associated with the job; and

wherein the plurality of instructions, when executed, further cause the compute device to associate the tasks within the job with a workload phase identifier.

19. The one or more non-transitory machine-readable storage media of claim 13 , wherein to send an availability request to a micro-orchestrator comprises to send the availability request to multiple micro-orchestrators.

20. The one or more non-transitory machine-readable storage media of claim 13 , wherein to receive availability data from the micro-orchestrator comprises to receive an indication of an estimated time to complete the tasks accepted by a micro-orchestrator.

21. The one or more non-transitory machine-readable storage media of claim 13 , wherein to receive availability data from the micro-orchestrator comprises to receive an indication of whether an accelerator sled associated with the micro-orchestrator can access a shared memory with another accelerator sled for parallel execution of one or more of the tasks.

22. The one or more non-transitory machine-readable storage media of claim 13 , wherein the accelerator sled is one of a plurality of accelerator sleds and wherein to assign the tasks to the micro-orchestrator comprises to assign the tasks as a function of a best fit of each associated accelerated sled to the tasks.

23. The one or more non-transitory machine-readable storage media of claim 22 , wherein to assign the tasks as a function of a best fit of each associated accelerated sled comprises to consolidate tasks on associated accelerator sleds to reduce network congestion.

24. The one or more non-transitory machine-readable storage media of claim 22 , wherein to assign the tasks as a function of a best fit of each associated accelerated sled comprises to assign the tasks as a function of estimated time completion of each task.

25. A compute device comprising:

circuitry for receiving a request from a compute sled to accelerate execution of a job, wherein the job includes a set of tasks;

means for analyzing the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks;

circuitry for sending an availability request to a micro-orchestrator of an accelerator sled communicatively coupled to the compute device, wherein the availability request includes the metadata;

circuitry for receiving availability data from the micro-orchestrator, wherein the availability data is indicative of which of the tasks the micro-orchestrator has accepted for acceleration on an associated accelerator sled; and

means for assigning the tasks to the micro-orchestrator as a function of the availability data.

26. A method comprising:

receiving, by a compute device, a request from a compute sled to accelerate execution of a job, wherein the job includes a set of tasks;

analyzing, by the compute device, the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks;

sending, by the compute device, an availability request to a micro-orchestrator of an accelerator sled communicatively coupled to the compute device, wherein the availability request includes the metadata;

receiving, by the compute device, availability data from the micro-orchestrator, wherein the availability data is indicative of which of the tasks the micro-orchestrator has accepted for acceleration on an associated accelerator sled; and

assigning, by the compute device, the tasks to the micro-orchestrator as a function of the availability data.

27. The method of claim 26 , wherein receiving a request to accelerate a job comprises receiving a request that includes code indicative operations to be performed within the job; and

wherein analyzing the request comprises analyzing the code to identify operations to be grouped into tasks.

28. The method of claim 26 , wherein analyzing the request comprises determining a type of acceleration for each task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2017
From: BALLE, SUSANNE M.; BERNAT, FRANCESC GUIM; PUTYRSKI, SLAWOMIR; GRECCO, JOE; MITCHEL, HENRY; KHANNA, RAHUL; CUSTODIO, EVAN
To: INTEL CORPORATION
Reel/Frame 044154/0170 →
Priority Claims (1)
IN 201741030632 · Aug 30, 2017 · national
Continuity (2)
Provisional Application 62427268 · Nov 29, 2016
Related Publication 20180150298A1 · May 31, 2018
Cited By (5)
US 12,260,257 US 12,261,940 US 12,288,101 US 12,506,817 US 12,619,465