IP Library Granted Patent US 12,705,103
Granted Patent B2
US 12,705,103 · App. 18/355,351 · Granted Aug 11, 2026

Edge domain-specific accelerator virtualization and scheduling

Inventors: William Jeffery White (Plano, TX); Said Tabet (Austin, TX)
Assignee: DELL PRODUCTS L.P.
G06F9/5038G06F9/4881G06F9/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,103
App. No.
18/355,351
Granted
Aug 11, 2026
Kind
B2
Abstract

Presented herein are embodiments to implement a temporal queueing system with class-based fair queuing and dynamic resource allocation based on a novel look-ahead capability to manage various models and workloads for utilization/efficiency improvements. Embodiments may be implemented to allocate accelerator resources based on platform-defined timeslots, and therefore significantly increase the ability of workloads to access hardware accelerator resources. Training and inference may be supported with flexible preemption and the ability to support run-to-completion for training tasks while still supporting non-run-to-completion for inference tasks. Embodiments may be implemented by an edge software operation platform through virtual accelerators to allow emulation of different types of hardware accelerators and to map to the hardware accelerators with hardware-specific procedures managed by an edge orchestrator and an edge endpoint. Accordingly, embodiments of the present disclosure reduce the requirements for the workload to manage platform capacity and hardware.

Claims (44)

1 . A processor-implemented method for edge domain-specific accelerator (DSA) virtualization and scheduling comprising:

configuring, by an edge orchestrator (EO), a virtual accelerator to virtualize one or more DSAs in an edge endpoint;

associating the virtual accelerator with an application to be executed at the edge endpoint; and

assigning, by the EO, one or more resource utilization parameters to the virtual accelerator for executing the application at the edge endpoint using allocated DSA resources in a timeslot scheduled using a time division queuing.

2 . The processor-implemented method of claim 1 wherein the virtual accelerator virtualizes the one or more DSAs into a single resource pool at the edge endpoint.

3 . The processor-implemented method of claim 1 wherein the timeslot is allocated in the time division queuing with elastic dynamic allocation.

4 . The processor-implemented method of claim 3 wherein the time division queuing is realized as a class-based weighted fair queuing (CBWFQ) with strict priority queuing (SPQ).

5 . The processor-implemented method of claim 1 wherein the one or more resource utilization parameters comprise one or more of:

a minimum resource utilization required for task execution;

a maximum streaming multiprocessor (SM) utilization limit; and

a mean resource utilization as a target average utilization.

6 . The processor-implemented method of claim 1 wherein the timeslot is obtained based at least on a mean normalized accelerator unit (NAU) from a resource normalization framework.

7 . The processor-implemented method of claim 1 wherein the allocated DSA resources are determined based on one or more of:

a priority specified by a customer through a service plan/manifest;

a power consumption specific to edge deployment;

a cost for cloud domains;

an accelerator resource requirement estimated from an application resource uncertainty estimation process;

parameters related to streaming multiprocessor (SM)/logic block (LB) execution in real-time;

a task/job category as run-to-completion (RTC) or non-RTC (NRTC); and

a task/job category as preemptable or non-preemptable.

8 . A processor-implemented method for edge domain-specific accelerator (DSA) virtualization and scheduling comprising:

given an application to be executed in an edge endpoint, the application being associated with a virtual accelerator that is configured to virtualize one or more physical accelerators in the edge endpoint into a single resource pool:

implementing, by a queuing scheduler, a temporal queuing with time slicing to allocate the virtual accelerator a timeslot;

during the allocated timeslot, loading one or more models and workloads for the application into a memory of the physical accelerators for application execution; and

executing the application in the allocated timeslot using the one or more models and workloads.

9 . The processor-implemented method of claim 8 wherein the timeslot is allocated using a time division queuing with elastic dynamic allocation.

10 . The processor-implemented method of claim 9 wherein the time division queuing is realized as a class-based weighted fair queuing (CBWFQ) with strict priority queuing (SPQ).

11 . The processor-implemented method of claim 10 wherein timeslot allocation for a class is managed dynamically based on one or more of:

a category of job/task of the class as run-to-completion (RTC) or non-RTC (NRTC); and

a category of job/task of the class as preemptable or non-preemptable.

12 . The processor-implemented method of claim 8 wherein the allocated timeslot is obtained based at least on a mean normalized accelerator unit (NAU) from a resource normalization framework.

13 . The processor-implemented method of claim 8 further comprising:

removing the one or more models from the memory by an end of the allocated timeslot such that the physical accelerators are ready for executing another application at a next timeslot.

14 . The processor-implemented method of claim 8 further comprising:

responsive to the virtual accelerator not being able to submit all data to be completed during the allocated timeslot, queuing the workload at the virtual accelerator until a next timeslot cycle.

15 . A non-transitory computer-readable medium or media comprising one or more sequences of instructions which, when executed by at least one processor, cause steps to be performed comprising:

configuring a virtual accelerator to virtualize one or more domain-specific accelerators (DSAs) in an edge endpoint;

associating the virtual accelerator with an application to be executed at the edge endpoint; and

assigning one or more resource utilization parameters to the virtual accelerator for executing the application at the edge endpoint using allocated DSA resources in a timeslot scheduled using a time division queuing.

16 . The non-transitory computer-readable medium or media of claim 15 wherein the one or more DSAs are virtualized into a single resource pool at the edge endpoint.

17 . The non-transitory computer-readable medium or media of claim 15 wherein the timeslot is allocated in the time division queuing with elastic dynamic allocation.

18 . The non-transitory computer-readable medium or media of claim 17 wherein the time division queuing is realized as a class-based weighted fair queuing (CBWFQ) with strict priority queuing (SPQ).

19 . The non-transitory computer-readable medium or media of claim 15 wherein timeslot allocation for the application is based at least on a category of the application as a run-to-completion (RTC) task, which is generally non-preemptable, or a non-RTC (NRTC) task, which is generally preemptable.

20 . The non-transitory computer-readable medium or media of claim 15 wherein the timeslot is obtained based at least on a mean normalized accelerator unit (NAU) from a resource normalization framework.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2024
From: WHITE, WILLIAM JEFFERY; TABET, SAID
To: DELL PRODUCTS L.P.
Reel/Frame 067157/0860 →
Continuity (2)
Provisional Application 63450237 · Mar 6, 2023
Related Publication 20240303124A1 · Sep 12, 2024
References Cited (73)
US 8386495B1 · Sandler et al. · 2013 [cited by applicant]
US 10169101B2 · Banerjee · 2019 [cited by examiner]
US 10698717B2 · Tang · 2020 [cited by examiner]
US 10977078B2 · Rehman · 2021 [cited by applicant]
US 11171831B2 · Patel · 2021 [cited by examiner]
US 11228527B2 · Bangalore Krishnamurthy · 2022 [cited by applicant]
US 11356349B2 · Cui et al. · 2022 [cited by applicant]
US 11836656B2 · Cai et al. · 2023 [cited by applicant]
US 11966788B2 · MacDonald et al. · 2024 [cited by applicant]
US 12250159B2 · Chaurasia · 2025 [cited by examiner]
US 20040103387A1 · Teig et al. · 2004 [cited by applicant]
US 20230138568A1 · Singh · 2023 [cited by applicant]
US 20230185472A1 · Higginson et al. · 2023 [cited by applicant]
US 20230244537A1 · Wang · 2023 [cited by applicant]
US 20230409871A1 · Xu et al. · 2023 [cited by applicant]
US 20240095090A1 · Saito · 2024 [cited by examiner]
US 20240205165A1 · Smith et al. · 2024 [cited by applicant]
US 20240220639A1 · Sahu · 2024 [cited by examiner]
US 20240259879A1 · Ranganath · 2024 [cited by examiner]
US 20240303121A1 · White et al. · 2024 [cited by applicant]
US 20240303124A1 · White et al. · 2024 [cited by applicant]
US 20240303127A1 · White et al. · 2024 [cited by applicant]
US 20240303128A1 · White et al. · 2024 [cited by applicant]
US 20240303129A1 · White et al. · 2024 [cited by applicant]
US 20240303130A1 · White et al. · 2024 [cited by applicant]
US 20240303134A1 · White et al. · 2024 [cited by applicant]
US 20240305535A1 · White et al. · 2024 [cited by applicant]
Non-Final Office Action (2676), including List of Ref. Cited by Examiner and Considered by Examiner, dated Jan. 23, 2026, in U.S. Appl. No. 18/366,507 (24 pgs). [cited by applicant]
Non-Final Office Action (2677), including List of Ref. Cited by Examiner and Considered by Examiner, dated Jan. 23, 2026, in U.S. Appl. No. 18/366,520 (23 pgs). [cited by applicant]
Non-Final Office Action (2678), including List of Ref. Cited by Examiner and Considered by Examiner, dated Jan. 23, 2026, in U.S. Appl. No. 18/366,538 (23 pgs). [cited by applicant]
Response to Non-Final Office Action, filed Jan. 25, 2026, U.S. Appl. No. 18/366,507. (18 pgs). [cited by applicant]
Response to Non-Final Office Action, filed Jan. 25, 2026, U.S. Appl. No. 18/366,520. (16 pgs). [cited by applicant]
Response to Non-Final Office Action, filed Jan. 25, 2026, U.S. Appl. No. 18/366,538. (17 pgs). [cited by applicant]
Kolosov, Oleg, et al. “Benchmarking in the Dark: On the Absence of Comprehensive Edge Datasets.” Proceedings of the 2nd USENIX Workshop on Hot Topics in Edge Computing, No date. [cited by applicant]
HotEdge), 2020. https://www.usenix.org/conference/hotedge20/presentation/kolosov. (11 pages). [cited by applicant]
Wiki contributors. “Jensen-Shannon divergence.” Wikipedia, The Free Encyclopedia, Mar. 21, 2025, https://en.wikipedia.org/wiki/Jensen%E2%80%93Shannon_divergence. [cited by applicant]
Accessed Mar. 21, 2025. (6 pages). [cited by applicant]
Salem, Osman, Farid Naït-Abdesselam, and Ahmed Mehaoua. “Anomaly Detection in Network Traffic using Jensen-Shannon Divergence,” Proceedings of the IEEE International, No date. [cited by applicant]
Conference on Communications (ICC), 2012, pp. 5200-5204. IEEE. https://doi.org/10.1109/ICC.2012.6364602. (6 pages). [cited by applicant]
Soos, Gabor, Daniel Ficzere, and Pal Varga. “Towards Traffic Identification and Modeling for 5G Application Use-Cases.” Electronics, vol. 9, No. 4, 2020, p. 640. [cited by applicant]
https://doi.org/10.3390/electronics9040640. (32 pages), No date. [cited by applicant]
Sisworo. “On Holder Exponents.” Jurnal Matematika dan Sains (JMS), vol. 4, No. 3, 1999, pp. 244-259. https://www.researchgate.net/publication/309421576_On_Holder_Exponents. [cited by applicant]
https://www.researchgate.net/publication/309421576_On_Holder_Exponents. (34 pages), No date. [cited by applicant]
Feng, Yihui, et al. “Scaling Large Production Clusters with Partitioned Synchronization.” Proceedings of the 2021 USENIX Annual Technical Conference (USENIX ATC '21). [cited by applicant]
Jul. 14-16, 2021. https://www.usenix.org/conference/atc21/presentation/feng-yihui. (16 pages). [cited by applicant]
Xie, Junfei, et al. “M-PCM-OFFD: An effective output statistics estimation method for systems of high dimensional uncertainties subject to low-order parameter interactions.” No date. [cited by applicant]
Mathematics and Computers in Simulation, vol. 159, 2019, pp. 93-118. https://doi.org/10.1016/j.matcom.2018.10.010. (26 pages). [cited by applicant]
Wiki contributors. “Erlang distribution.” Wikipedia, The Free Encyclopedia, Mar. 21, 2025, https://en.wikipedia.org/wiki/Erlang_distribution. Accessed Mar. 21, 2025. (6 pages). [cited by applicant]
Toczé, Klervie, et al. “Edge Workload Trace Gathering and Analysis for Benchmarking.” Proceedings of the 2022 IEEE 6th International Conference on Fog and Edge Computing. [cited by applicant]
ICFEC), 2022, pp. 34-41. https://doi.org/10.1109/ICFEC54809.2022.00012. (8 pages). [cited by applicant]
Qiu, Haoran, et al. “FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented Microservices.” Proceedings of the 14th USENIX Symposium on Operating, No date. [cited by applicant]
Systems Design and Implementation (OSDI), 2020, pp. 805-825. https://www.usenix.org/conference/osdi20/presentation/qiu. (22 pages). [cited by applicant]
Wang et al. “The Cost of Cloud, a Trillion Dollar Paradox.” Andreessen Horowitz, May 27, 2021, https://a16z.com/the-cost-of-cloud-a-trillion-dollar-paradox/ (12 pages). [cited by applicant]
Non-Final Office Action (2674), including List of Ref. Cited by Examiner and Considered by Examiner, dated Dec. 12, 2025, in U.S. Appl. No. 18/366,461 (22 pgs). [cited by applicant]
Response to Non-Final Office Action, filed Dec. 14, 2025, U.S. Appl. No. 18/366,461. (15 pgs). [cited by applicant]
Notice of Allowance mailed Sep. 3, 2025 for U.S. Appl. No. 18/366,490, 20 pages. [cited by applicant]
Liu, M., Wan, Y., Lin, Z., Lewis, F.L., Xie, J., Jalaian, B.A. (2021). Computational Intelligence in Uncertainty Quantification for Learning Control and Differential Games. [cited by applicant]
In: Vamvoudakis, K.G., Wan, Y., Lewis, F.L., Cansever, D. (eds) Handbook of Reinforcement Learning and Control. Studies in Systems, Decision and Control, vol. 325. Springer, No date. [cited by applicant]
Cham. https://doi.org/10.1007/978-3-030-60990-0_13 (34 pages ), No date. [cited by applicant]
Ghorbani, Amir, Yifan Wang, Yanzhi Xue, Massoud Pedram, and Paul Bogdan. “Prediction and Control of Bursty Cloud Workloads: A Fractal Framework.” University of Southern, No date. [cited by applicant]
California, 2014. https://dx.doi.org/10.1145/2656075.2656095 (9 pages). [cited by applicant]
Ali-Eldin, Ahmed, et al. “The Hidden Cost of the Edge: A Performance Comparison of Edge and Cloud Latencies.” Proceedings of the International Conference for High Performance, No date. [cited by applicant]
Computing, Networking, Storage and Analysis (SC '21), 2021, https://doi.org/10.1145/3458817.3476142 (15 pages). [cited by applicant]
Tirmazi, Muhammad, et al. “Borg: the Next Generation.” Proceedings of the Fifteenth European Conference on Computer Systems (EuroSys '20), Apr. 27-30, 2020, Heraklion, Greece. [cited by applicant]
ACM, New York, NY, USA, 2020, pp. 1-14. https://doi.org/10.1145/3342195.3387517 (14 pages). [cited by applicant]
Notice of Allowance, including References Considered by Examiner, mailed Mar. 13, 2026 for U.S. Appl. No. 18/366,507, 15 pages. [cited by applicant]
Notice of Allowance, including References Considered by Examiner, mailed Mar. 13, 2026 for U.S. Appl. No. 18/366,520, 20 pages. [cited by applicant]
Notice of Allowance, including References Considered by Examiner, mailed Mar. 13, 2026 for U.S. Appl. No. 18/366,538, 20 pages. [cited by applicant]
Supplemental Notice of Allowance (3rd) mailed Mar. 10, 2026 for U.S. Appl. No. 18/366,461, 2 pages. [cited by applicant]
Notice of Allowance (2nd) mailed Feb. 13, 2025 for U.S. Appl. No. 18/366,490, 10 pages. [cited by applicant]
Notice of Allowance mailed Feb. 13, 2026 for U.S. Appl. No. 18/366,549, 32 pages. [cited by applicant]
Notice of Allowance mailed Feb. 19, 2026 for U.S. Appl. No. 18/366,555, 52 pages. [cited by applicant]
Notice of Allowance mailed Feb. 3, 2026 for U.S. Appl. No. 18/366,461, 7 pages. [cited by applicant]