IP Library › Granted Patent US 12,314,778
Granted Patent B2
US 12,314,778 · App. 17/482,962 · Granted May 27, 2025

Systems and methods for data processing unit aware workload migration in a virtualized datacenter environment

Inventors: John Kelly (Mallow, IE); Dharmesh M. Patel (Round Rock, TX)
Assignee: DELL PRODUCTS L.P.
G06F9/5088G06F9/455G06F9/45533G06F9/48G06F9/4806G06F9/4843G06F9/485G06F9/4856G06F9/4881G06F9/50G06F9/5027G06F9/505G06F9/5061G06F9/5077G06F9/5083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,778
App. No.
17/482,962
Filed
Sep 23, 2021
Granted
May 27, 2025
Kind
B2
Art Unit
2196
USPC
718/104
Abstract

Techniques described herein relate to systems and methods for data processing unit (DPU) workload management. Such methods may include obtaining, by a DPU workload manager, a DPU utilization value for a first node in a device ecosystem; making a first determination, by the DPU workload manager and using the DPU utilization value, that DPU utilization of the first node is above a DPU utilization threshold configured for the first node; identifying, in response to the first determination, a workload executing on the first node as a migration candidate based at least in part on a central processing unit (CPU) utilization value associated with the workload; initiating a migration of the migration candidate to a second node in the device ecosystem; and obtaining, after the migration completes, a second DPU utilization value for the first node to determine whether the second DPU utilization value is below the DPU utilization threshold.

Claims (54)

1. A method for data processing unit (DPU) workload management, the method comprising:

obtaining, by a DPU workload manager, a DPU utilization value for a first node in a device ecosystem, wherein the DPU utilization value comprises a percentage of total resources being utilized by a central processing unit (CPU), a network interface, and a logic device;

making a first determination, by the DPU workload manager and using the DPU utilization value, that DPU utilization of the first node is above a DPU utilization threshold configured for the first node;

identifying, in response to the first determination, a workload executing on the first node as a migration candidate, wherein:

the workload is one of a plurality of workloads executing on the first node, and identifying the workload as the migration candidate comprises:

obtaining a CPU utilization information set associated with the first node, wherein the CPU utilization information set includes overall CPU utilization of the first node;

obtaining a DPU utilization information set associated with the first node, wherein the DPU utilization information set includes overall DPU utilization of the first node;

calculating, using the CPU utilization information set and the DPU utilization information set, a plurality of correlation coefficients each corresponding to one of the plurality of workloads; and

selecting the workload, from among the plurality of workloads, with a highest corresponding correlation coefficient;

initiating a migration of the migration candidate to a second node in the device ecosystem, wherein the second node is a migration destination for the migration candidate based on the second node having a lowest DPU utilization value among a plurality of nodes of the device ecosystem and having appropriate resources to handle operational requirements of the migration candidate;

obtaining, after the migration completes, a second DPU utilization value for the first node; and

making a second determination, after the obtaining of the second DPU utilization value, that the second DPU utilization value is below the DPU utilization threshold.

2. The method of claim 1 , wherein the plurality of correlation coefficients are Pearson correlation coefficients.

3. The method of claim 1 , wherein when the second DPU utilization value is not below the DPU utilization threshold, the method further comprises:

identifying a second workload executing on the first node as a second migration candidate based at least in part on a second CPU utilization value associated with the second workload; and

initiating a second migration of the second migration candidate to a third node in the device ecosystem.

4. The method of claim 1 , wherein the second node is a migration destination for the migration candidate based on the second node having a second node DPU utilization value below a second node DPU utilization threshold and having a second node CPU utilization value below a second node CPU utilization threshold.

5. A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for data processing unit (DPU) workload management, the method comprising:

obtaining, by a DPU workload manager, a DPU utilization value for a first node in a device ecosystem, wherein the DPU utilization value comprises a percentage of total resources being utilized by a central processing unit (CPU), a network interface, and a logic device;

making a first determination, by the DPU workload manager and using the DPU utilization value, that DPU utilization of the first node is above a DPU utilization threshold configured for the first node;

identifying, in response to the first determination, a workload executing on the first node as a migration candidate, wherein:

the workload is one of a plurality of workloads executing on the first node, and identifying the workload as the migration candidate comprises:

obtaining a CPU utilization information set associated with the first node, wherein the CPU utilization information set includes overall CPU utilization of the first node;

obtaining a DPU utilization information set associated with the first node, wherein the DPU utilization information set includes overall DPU utilization of the first node;

calculating, using the CPU utilization information set and the DPU utilization information set, a plurality of correlation coefficients each corresponding to one of the plurality of workloads; and

selecting the workload, from among the plurality of workloads, with a highest corresponding correlation coefficient;

initiating a migration of the migration candidate to a second node in the device ecosystem, wherein the second node is a migration destination for the migration candidate based on the second node having a lowest DPU utilization value among a plurality of nodes of the device ecosystem and having appropriate resources to handle operational requirements of the migration candidate;

obtaining, after the migration completes, a second DPU utilization value for the first node; and

making a second determination, after the obtaining of the second DPU utilization value, that the second DPU utilization value is below the DPU utilization threshold.

6. The non-transitory computer readable medium of claim 5 , wherein the plurality of correlation coefficients are Pearson correlation coefficients.

7. The non-transitory computer readable medium of claim 5 , wherein when the second DPU utilization value is not below the DPU utilization threshold, the method further comprises:

identifying a second workload executing on the first node as a second migration candidate based at least in part on a second CPU utilization value associated with the second workload; and

initiating a second migration of the second migration candidate to a third node in the device ecosystem.

8. The non-transitory computer readable medium of claim 5 , wherein the second node is a migration destination for the migration candidate based on the second node having a second node DPU utilization value below a second node DPU utilization threshold and having a second node CPU utilization value below a second node CPU utilization threshold.

9. A system for data processing unit (DPU) workload management, the system comprising:

a processor comprising circuitry;

memory; and

a DPU workload manager operatively connected to a plurality of nodes of a device ecosystem, executing on the processor and using the memory, and configured to:

obtain a DPU utilization value for a first node of the plurality of nodes, wherein the DPU utilization value comprises a percentage of total resources being utilized by a central processing unit (CPU), a network interface, and a logic device;

make a first determination, using the DPU utilization value, that DPU utilization of the first node is above a DPU utilization threshold configured for the first node;

identify, in response to the first determination, a workload executing on the node as a migration candidate, wherein:

the workload is one of a plurality of workloads executing on the first node, and

identifying the workload as the migration candidate comprises:

obtaining a CPU utilization information set associated with the first node, wherein the CPU utilization information set includes overall CPU utilization of the first node;

obtaining a DPU utilization information set associated with the first node, wherein the DPU utilization information set includes overall DPU utilization of the first node;

calculating, using the CPU utilization information set and the DPU utilization information set, a plurality of correlation coefficients each corresponding to one of the plurality of workloads; and

selecting the workload, from among the plurality of workloads, with a highest corresponding correlation coefficient;

initiate a migration of the migration candidate to a second node in the device ecosystem, wherein the second node is a migration destination for the migration candidate based on the second node having a lowest DPU utilization value among a plurality of nodes of the device ecosystem and having appropriate resources to handle operational requirements of the migration candidate;

obtain, after the migration completes, a second DPU utilization value for the first node; and

making a second determination, after the obtaining of the second DPU utilization value, that the second DPU utilization value is below the DPU utilization threshold.

10. The system of claim 9 , wherein when the second DPU utilization value is not below the DPU utilization threshold, the DPU workload manager is further configured to:

identify a second workload executing on the first node as a second migration candidate based at least in part on a second central processing unit CPU utilization value associated with the second workload; and

initiate a second migration of the second migration candidate to a third node in the device ecosystem.

11. The system of claim 9 , wherein the second node is a migration destination for the migration candidate based on the second node having a second node DPU utilization value below a second node DPU utilization threshold and having a second node CPU utilization value below a second node CPU utilization threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2021
From: KELLY, JOHN; PATEL, DHARMESH M.
To: DELL PRODUCTS L.P.
Reel/Frame 057592/0115 →
Continuity (1)
Related Publication 20230091753A1 · Mar 23, 2023
References Cited (12)
US 10684888B1 · Sethuramalingam · 2020 [cited by examiner]
US 11307885B1 · Luciano · 2022 [cited by examiner]
US 11886926B1 · Gadalin · 2024 [cited by examiner]
US 20090228589A1 · Korupolu · 2009 [cited by examiner]
US 20120254414A1 · Scarpelli · 2012 [cited by examiner]
US 20140181834A1 · Lim · 2014 [cited by examiner]
US 20150347183A1 · Borthakur · 2015 [cited by examiner]
US 20170315838A1 · Nidugala · 2017 [cited by examiner]
US 20190065281A1 · Bernat · 2019 [cited by examiner]
US 20200366733A1 · Parvataneni · 2020 [cited by examiner]
US 20220046036A1 · Bastawala · 2022 [cited by examiner]
US 20230031998A1 · Boyapalle · 2023 [cited by examiner]