IP Library › Granted Patent US 12,260,257
Granted Patent B2
US 12,260,257 · App. 17/214,605 · Granted Mar 25, 2025

Technologies for offloading acceleration task scheduling operations to accelerator sleds

Inventors: Susanne M. Balle (Hudson, NH); Francesc Guim Bernat (Barcelona, ES); Slawomir Putyrski (Gdynia, PL); Joe Grecco (Saddle Brook, NJ); Henry Mitchel (Wayne, NJ); Rahul Khanna (Portland, OR); Evan Custodio (North Attleboro, MA)
Assignee: Intel Corporation
G06F9/505G06F3/0604G06F3/0608G06F3/0611G06F3/0613G06F3/0617G06F3/0641G06F3/0647G06F3/065G06F3/0653G06F3/067G06F7/06G06F8/65G06F8/654G06F8/656G06F8/658G06F9/3891G06F9/4401G06F9/45533G06F9/4843G06F9/4881G06F9/5005G06F9/5038G06F9/5044G06F9/5083G06F9/544G06F11/0709G06F11/0751G06F11/079G06F11/3006G06F11/3034G06F11/3055G06F11/3079G06F11/3409G06F12/0284G06F12/0692G06F13/1652G06F16/1744G06F21/57G06F21/6218G06F21/73G06F21/76G06T1/20G06T1/60G06T9/005H01R13/453H01R13/4536H01R13/4538H01R13/631H03K19/1731H03M7/3084H03M7/40H03M7/42H03M7/60H03M7/6011H03M7/6017H03M7/6029H04L9/0822H04L12/2881H04L12/4633H04L41/044H04L41/0816H04L41/0853H04L41/12H04L43/04H04L43/06H04L43/08H04L43/0894H04L47/20H04L47/2441H04L49/104H04L61/5007H04L67/10H04L67/1014H04L67/63H04L67/75H05K7/1452H05K7/1487H05K7/1491G06F11/1453G06F12/023G06F15/80G06F16/285G06F2212/401G06F2212/402G06F2221/2107H04L41/046H04L41/0896H04L41/142H04L47/78H04L63/1425H04Q11/0005H05K7/1447H05K7/1492
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,257
App. No.
17/214,605
Granted
Mar 25, 2025
Kind
B2
Abstract

Technologies for offloading acceleration task scheduling operations to accelerator sleds include a compute device to receive a request from a compute sled to accelerate the execution of a job, which includes a set of tasks. The compute device is also to analyze the request to generate metadata indicative of the tasks within the job, a type of acceleration associated with each task, and a data dependency between the tasks.

Claims (30)

1. One or more non-transitory machine-readable storage media comprising a plurality of instructions stored thereon that, in response to being executed, cause one or more processors to:

execute an orchestrator to:

receive a request to perform a workload, wherein the workload includes a set of tasks;

determine availability data of multiple accelerator devices to perform the workload, wherein the availability data is indicative of which of the one or more of the set of tasks are acceptable for execution on an associated accelerator device of the multiple accelerator devices and which of the one or more of the set of tasks are not acceptable for execution on an associated accelerator device of the multiple accelerator devices;

select one or more of the multiple accelerator devices to perform one or more of the set of tasks based on the availability data and also based on capability to complete the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on a type of acceleration for the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on capability for parallel execution of the one or more of the set of tasks, and wherein the type of acceleration for the one or more of the set of tasks comprises one or more of: cryptographic operation or data compression; and

cause execution of the one or more of the set of tasks on the selected one or more of the multiple accelerator devices.

2. The one or more non-transitory machine-readable storage media of claim 1 , wherein the capability to complete the one or more of the set of tasks is based on data dependence between the one or more of the set of tasks.

3. The one or more non-transitory machine-readable storage media of claim 1 , wherein the capability to complete the one or more of the set of tasks is based on subdivision of the one or more of the set of tasks to operate on different portions of a data set concurrently.

4. The one or more non-transitory machine-readable storage media of claim 1 , wherein the capability to complete the one or more of the set of tasks is based on an estimated time to complete the one or more of the set of tasks.

5. The one or more non-transitory machine-readable storage media of claim 1 , wherein the capability to complete the one or more of the set of tasks is based on network congestion.

6. The one or more non-transitory machine-readable storage media of claim 1 , wherein the capability to complete the one or more of the set of tasks is based on an earliest received indication of capability to complete the one or more of the set of tasks.

7. The one or more non-transitory machine-readable storage media of claim 1 , wherein the one or more accelerator devices comprise one or more of: a processor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or a graphics processing unit (GPU).

8. An apparatus comprising:

an input/output (I/O) subsystem and

circuitry to:

receive a request to perform a workload, wherein the workload includes a set of tasks;

determine availability data of multiple accelerator devices to perform the workload, wherein the availability data is indicative of which of the one or more of the set of tasks are acceptable for execution on an associated accelerator device of the multiple accelerator devices and which of the one or more of the set of tasks are not acceptable for execution on an associated accelerator device of the multiple accelerator devices;

select one or more of the multiple accelerator devices to perform one or more of the set of tasks based on the availability data and also based on capability to complete the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on a type of acceleration for the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on capability for parallel execution of the one or more of the set of tasks, and wherein the type of acceleration for the one or more of the set of tasks comprises one or more of: cryptographic operation or data compression; and

cause execution of the one or more of the set of tasks on the selected one or more of the multiple accelerator devices.

9. The apparatus of claim 8 , wherein the capability to complete the one or more of the set of tasks is based on data dependence between the one or more of the set of tasks.

10. The apparatus of claim 8 , wherein the capability to complete the one or more of the set of tasks is based on subdivision of the one or more of the set of tasks to operate on different portions of a data set concurrently.

11. The apparatus of claim 8 , wherein the capability to complete the one or more of the set of tasks is based on an estimated time to complete the one or more of the set of tasks.

12. The apparatus of claim 8 , comprising the one or more accelerator devices coupled to the I/O subsystem.

13. A method comprising:

receiving a request to perform a workload, wherein the workload includes a set of tasks;

determining availability data of multiple accelerator devices to perform the workload, wherein the availability data is indicative of which of the one or more of the set of tasks are acceptable for execution on an associated accelerator device of the multiple accelerator devices and which of the one or more of the set of tasks are not acceptable for execution on an associated accelerator device of the multiple accelerator devices;

selecting one or more of the multiple accelerator devices to perform one or more of the set of tasks based on the availability data and also based on capability to complete the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on a type of acceleration for the one or more of the set of tasks, wherein the capability to complete the one or more of the set of tasks is based on capability for parallel execution of the one or more of the set of tasks, and wherein the type of acceleration for the one or more of the set of tasks comprises one or more of: cryptographic operation or data compression; and

causing execution of the one or more of the set of tasks on the selected one or more of the multiple accelerator devices.

14. The method of claim 13 , wherein the capability to complete the one or more of the set of tasks is based on subdivision of the one or more of the set of tasks to operate on different portions of a data set concurrently.

15. The method of claim 13 , wherein the capability to complete the one or more of the set of tasks is based on an estimated time to complete the one or more of the set of tasks.

Priority Claims (1)
IN 201741030632 · Aug 30, 2017 · national
Continuity (3)
Continuation 15721821 · Sep 30, 2017
Provisional Application 62427268 · Nov 29, 2016
Related Publication 20210318823A1 · Oct 14, 2021
References Cited (40)
US 7694160B2 · Esliger · 2010 [cited by examiner]
US 8082418B2 · Stillwell, Jr. · 2011 [cited by examiner]
US 8284844B2 · MacInnis · 2012 [cited by examiner]
US 8463865B2 · Seigneret · 2013 [cited by examiner]
US 8776066B2 · Krishnamurthy · 2014 [cited by examiner]
US 8914805B2 · Krishnamurthy · 2014 [cited by examiner]
US 9026765B1 · Marshak et al. · 2015 [cited by applicant]
US 9612651B2 · Luan · 2017 [cited by examiner]
US 9891935B2 · Teh · 2018 [cited by examiner]
US 9996394B2 · Rossbach · 2018 [cited by examiner]
US 10034407B2 · Miller et al. · 2018 [cited by applicant]
US 10045098B2 · Adiletta et al. · 2018 [cited by applicant]
US 10085358B2 · Adiletta et al. · 2018 [cited by applicant]
US 10558490B2 · Ronen · 2020 [cited by examiner]
US 10831547B2 · Suzuki · 2020 [cited by examiner]
US 10949362B2 · Balle · 2021 [cited by examiner]
US 10963176B2 · Balle · 2021 [cited by examiner]
US 10990309B2 · Bernat et al. · 2021 [cited by applicant]
US 11029870B2 · Balle et al. · 2021 [cited by applicant]
US 20100191823A1 · Archer et al. · 2010 [cited by applicant]
US 20110131580A1 · Krishnamurthy · 2011 [cited by examiner]
US 20120054770A1 · Krishnamurthy et al. · 2012 [cited by applicant]
US 20120054771A1 · Krishnamurthy · 2012 [cited by examiner]
US 20130179485A1 · Chapman et al. · 2013 [cited by applicant]
US 20130232495A1 · Rossbach et al. · 2013 [cited by applicant]
US 20140359629A1 · Ronen · 2014 [cited by examiner]
US 20150007182A1 · Rossbach et al. · 2015 [cited by applicant]
US 20170046179A1 · Teh et al. · 2017 [cited by applicant]
US 20170116004A1 · Devegowda et al. · 2017 [cited by applicant]
US 20170317945A1 · Guo et al. · 2017 [cited by applicant]
US 20180077235A1 · Nachimuthu et al. · 2018 [cited by applicant]
US 20190065281A1 · Bernat et al. · 2019 [cited by applicant]
CN 105979007A · 2016 [cited by applicant]
First Office Action for U.S. Appl. No. 17/221,541, Mailed Oct. 25, 2022, 15 pages. [cited by applicant]
Artail et al. “Speedy Cloud: Cloud Computing with Support for Hardware Acceleration Services”, 2017 IEEE, pp. 850-865. [cited by applicant]
Asiatici et al. “Virtualized Execution Runtime for FPGA Accelerators in the Cloud”, 2017 IEEE, pp. 1900-1910. [cited by applicant]
Caulfield et al, A Cloud-scale acceleration Architecture, Microsoft Corp, Oct. 2016, 13 pages. [cited by applicant]
Diamantopoulos et al. “High-level Synthesizable Dataflow Map Reduce Accelerator for FPGA-coupled Data Centers”, 2015 IEEE, pp. 1-8. [cited by applicant]
Ding et al. “A Unified OpenCL-flavor Programming Model with Scalable Hybrid Hardware Platform on FPGAs”, 2014 IEEE, 7 pages. [cited by applicant]
Fahmy et al. “Virtualized FPGA Accelerators for Efficient Cloud Computing”, 2015 IEEE, pp. 430-435. [cited by applicant]
Cited By (1)
US 12,561,057