IP Library › Granted Patent US 12,748,611
Granted Patent B2
US 12,748,611 · App. 17/977,942 · Granted Sep 29, 2026

Facilitating workload migration in data centers using virtual machine management

Inventors: Andrew Currid (Alameda, CA); Anshul Fadnavis (San Jose, CA); Chenghuan Jia (Fremont, CA); Ankit Agrawal (San Jose, CA)
Assignee: NVIDIA CORPORATION
G06F9/45558G06F2009/4557
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,611
App. No.
17/977,942
Granted
Sep 29, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to determine that a first group including first hardware components is compatible with a second group including second hardware components based at least on a first label associated with the first group and a second label associated with the second group, and cause at least one workload to be migrated from the first group to the second group based at least on determining the first and second groups are compatible with one another.

Claims (114)

1 . A method comprising:

encoding at least a portion of a first topology of first hardware components of a first group into at least one first encoded topology;

creating a first label associated with the first group using the at least one first encoded topology and identifications of the first hardware components;

encoding at least a portion of a second topology of second hardware components of a second group into at least one second encoded topology;

creating a second label associated with the second group using the at least one second encoded topology and identifications of the second hardware components;

determining the first label matches the second label; and

causing at least one workload to be migrated from the first group to the second group based at least on having determined that the first label matches the second label.

2 . The method of claim 1 , wherein the first label lists the identifications of the first hardware components, and the second label lists the identifications of the second hardware components in a same order.

3 . The method of claim 1 , further comprising:

causing a first portion of the at least one workload to be performed using a first virtual machine executed using the first hardware components, the second group being associated with an ordered list that lists the second hardware components in an order,

causing the second hardware components to be mapped to a second virtual machine in accordance with the order, and

causing a second portion of the at least one workload to be performed using the second virtual machine.

4 . The method of claim 3 , further comprising:

obtaining the ordered list; and

associating the ordered list with the second group,

wherein obtaining the ordered list comprises:

defining one or more switch groups for one or more switches in the second hardware components,

including as members of the one or more switch groups any of at least a portion of the second hardware components connected to the one or more switches, and

ordering at least a portion of the members of the one or more switch groups in accordance with a predefined order.

5 . The method of claim 4 , wherein obtaining the ordered list further comprises placing any of the second hardware components that is not a member of the one or more switch groups at a predefined location in the ordered list.

6 . The method of claim 4 , wherein obtaining the ordered list further comprises sorting the one or more switch groups within the ordered list.

7 . The method of claim 1 , wherein the at least one first encoded topology comprises a first topology string, and the at least one second encoded topology comprises a second topology string.

8 . The method of claim 1 , further comprising:

creating at least one first list of the identifications of the first hardware components, wherein the first label is created using the at least one first encoded topology and the at least one first list; and

creating at least one second list of the identifications of the second hardware components, wherein the second label is created using the at least one second encoded topology and the at least one second list.

9 . The method of claim 8 , wherein the at least one first list comprises a list of any graphics processing units in the first hardware components and the at least one second list comprises a list of any graphics processing units in the second hardware components.

10 . The method of claim 9 , wherein the at least one first list comprises a list of any of the first hardware components that is not a graphics processing unit and the at least one second list comprises a list of any of the second hardware components that is not a graphics processing unit.

11 . The method of claim 1 , wherein the first topology comprises a first portion defined at least in part by one or more first type connections, and a second portion defined at least in part by one or more second type connections,

the second topology comprises a third portion defined at least in part by the one or more first type connections, and a fourth portion defined at least in part by one or more second type connections,

the at least one first encoded topology comprises a first string encoding the first portion and a second string encoding the second portion, and

the at least one second encoded topology comprises a third string encoding the third portion and a fourth string encoding the fourth portion.

12 . The method of claim 1 , further comprising:

identifying the first group by determining first metrics for first paths connecting the first hardware components and selecting the first group based at least on the first metrics; and

identifying the second group by determining second metrics for second paths connecting the second hardware components and selecting the second group based at least on the second metrics.

13 . A system comprising:

one or more processors to:

encode at least a portion of a first topology of first hardware components of a first group into at least one first encoded topology;

create a first label associated with the first group using the at least one first encoded topology and identifications of the first hardware components;

encode at least a portion of a second topology of second hardware components of a second group into at least one second encoded topology;

create a second label associated with the second group using the at least one second encoded topology and identifications of the second hardware components;

determine the first label matches the second label; and

cause at least one workload to be migrated from the first group to the second group.

14 . The system of claim 13 , wherein the first label lists the identifications of the first hardware components, and the second label lists the identifications of the second hardware components in a same order.

15 . The system of claim 13 , further comprising:

a first computing device comprising the one or more processors;

a second computing device comprising the first group of hardware components; and

a third computing device comprising the second group of hardware components.

16 . The system of claim 13 , further comprising:

at least one first computing device comprising the one or more processors; and

at least one second computing device comprising the first group of hardware components and the second group of hardware components.

17 . The system of claim 13 , wherein the one or more processors are to:

cause a first portion of the at least one workload to be performed using a first virtual machine; and

cause a second virtual machine to be created on the second group, wherein causing the at least one workload to be migrated from the first group to the second group comprises causing the second group to be mapped to the second virtual machine in accordance with a hardware component list and causing a second portion of the at least one workload to be performed using the second virtual machine.

18 . The system of claim 17 , wherein the one or more processors are to obtain the hardware component list by:

defining one or more switch groups for one or more switches in the second hardware components, the one or more switch groups comprising, as members, any of at least a portion of the second hardware components connected to the one or more switches; and

ordering at least a portion of the members of the one or more switch groups.

19 . The system of claim 18 , wherein obtaining the hardware component list further comprises placing any of the second hardware components that is not a member of the one or more switch groups at a predefined location in the hardware component list.

20 . The system of claim 19 , wherein obtaining the hardware component list further comprises sorting the one or more switch groups within the hardware component list.

21 . The system of claim 13 , wherein the one or more processors are to:

create at least one first list of the identifications of the first hardware components, wherein the first label is created using the at least one first encoded topology and the at least one first list; and

create at least one second list of the identifications of the second hardware components, wherein the second label is created using the at least one second encoded topology and the at least one second list.

22 . The system of claim 21 , wherein the at least one first list comprises a list of any graphics processing units in the first hardware components and the at least one second list comprises a list of any graphics processing units in the second hardware components.

23 . The system of claim 22 , wherein the at least one first list comprises a list including any of the first hardware components other than graphics processing units and the at least one second list comprises a list including any of the second hardware components other than graphics processing units.

24 . The system of claim 21 , wherein the first topology comprises a first connection topology and a second connection topology, the first connection topology being defined at least in part by at least one first connection that connects hardware components in the portion of the first hardware components, the second connection topology being defined at least in part by at least one second connection that connects hardware components in the portion of the first hardware components, the at least one first and second connections having first and second connection types, respectively; and

the second topology comprises a third connection topology and a fourth connection topology, the third connection topology being defined at least in part by at least one third connection that connects hardware components in the portion of the second hardware components, the fourth connection topology being defined at least in part by at least one fourth connection that connects hardware components in the portion of the second hardware components, the at least one third and fourth connections having the first and second connection types, respectively.

25 . The system of claim 13 , wherein the one or more processors are to:

identify the first groups based at least on first weights associated with one or more first paths between at least a portion of the first hardware components; and

identify the second group based at least on second weights associated with one or more second paths between at least a portion of the second hardware components.

26 . The system of claim 13 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a first system for performing simulation operations;

a second system for performing deep learning operations;

a third system implemented using an edge device;

a fourth system implemented using a robot;

a fifth system incorporating one or more virtual machines (VMs);

a sixth system implemented at least partially in a data center;

a seventh system for performing digital twin operations;

an eighth system for performing light transport simulation;

a nineth system for performing collaborative content creation for 3D assets;

a tenth system for performing conversational Artificial Intelligence operations;

an eleventh system for generating synthetic data;

a twelfth system for implementing a web-hosted service for detecting program workload inefficiencies;

an application as an application programming interface (“API”);

a thirteenth system implemented at least partially using cloud computing resources; or

a fourteenth system for presenting one or more of virtual reality content, augmented reality content, or mixed reality content.

27 . A processor comprising:

one or more circuits to:

obtain at least one first encoded topology encoding at least a portion of a first topology of first hardware components;

generate a first label associated with the first hardware components using the at least one first encoded topology and identifications of the first hardware components;

obtain at least one second encoded topology encoding at least a portion of a second topology of second hardware components;

generate a second label associated with the second hardware components using the at least one second encoded topology and identifications of the second hardware components;

cause the first hardware components to stop performing at least one workload before the at least one workload is finished;

determine the second label matches the first label; and

cause the second hardware components to resume performing the at least one workload using state information obtained from the first hardware components.

28 . The processor of claim 27 , wherein the one or more circuits are to:

identify the first hardware components based at least on first weights associated with one or more first paths between at least a portion of the first hardware components; and

identify the second hardware components based at least on second weights associated with one or more second paths between at least a portion of the second hardware components.

29 . The processor of claim 27 , wherein the first label lists the identifications of the first hardware components, and the second label lists the identifications of the second hardware components in a same order.

30 . The processor of claim 27 , wherein the first and second topologies comprise first and second physical topologies, respectively, defined by first type connections,

the first and second topologies comprise third and fourth physical topologies, respectively, defined by second type connections,

the at least one first encoded topology encodes the first physical topology and the third physical topology,

the first label encodes the identifications of the first hardware components, the first physical topology, and the third physical topology,

the at least one second encoded topology encodes the second physical topology and the fourth physical topology, and

the second label encodes the identifications of the second hardware components, the second physical topology, and the fourth physical topology.

31 . The processor of claim 27 , wherein before the first hardware components stopped performing the at least one workload, the at least one workload was performed by a first virtual machine executing on the first hardware components, and

the one or more circuits are to map the second hardware components to a second virtual machine in a specified order and cause the second virtual machine to resume performing the at least one workload.

32 . The processor of claim 31 , wherein the one or more circuits are to:

determine the specified order by organizing at least a portion of the second hardware components into one or more switch groups, and ordering the one or more switch groups in accordance with a predefined order.

33 . The processor of claim 32 , wherein determining the specified order further comprises:

placing any of the second hardware components that is not a member of the one or more switch groups at a predefined location in the specified order.

34 . The processor of claim 27 , wherein the one or more circuits are to:

identify the first hardware components by determining at least one first expected path performance for at least one first path between the first hardware components, and selecting the first hardware components to begin performing the at least one workload based at least in part on the at least one first expected path performance; and

identify the second hardware components by determining at least one second expected path performance for at least one second path between the second hardware components, and selecting the second hardware components to resume performing the at least one workload based at least in part on the at least one second expected path performance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2022
From: CURRID, ANDREW; FADNAVIS, ANSHUL; JIA, CHENGHUAN; AGRAWAL, ANKIT
To: NVIDIA CORPORATION
Reel/Frame 061659/0190 →
Continuity (1)
Related Publication 20240143372A1 · May 2, 2024
References Cited (42)
US 8484654B2 · Graham · 2013 [cited by examiner]
US 8613085B2 · Diab et al. · 2013 [cited by applicant]
US 10026143B2 · Shu et al. · 2018 [cited by applicant]
US 10621001B1 · Braverman et al. · 2020 [cited by applicant]
US 11169883B1 · Burgin et al. · 2021 [cited by applicant]
US 20030204770A1 · Bergsten · 2003 [cited by applicant]
US 20060045092A1 · Kubsch et al. · 2006 [cited by applicant]
US 20100257269A1 · Clark · 2010 [cited by applicant]
US 20110066786A1 · Colbert · 2011 [cited by applicant]
US 20120144391A1 · Ueda · 2012 [cited by applicant]
US 20120216183A1 · Mahajan et al. · 2012 [cited by applicant]
US 20120243795A1 · Head et al. · 2012 [cited by applicant]
US 20120297236A1 · Ziskind et al. · 2012 [cited by applicant]
US 20120300669A1 · Zahavi · 2012 [cited by applicant]
US 20120324071A1 · Gulati et al. · 2012 [cited by applicant]
US 20130070515A1 · Mayhew et al. · 2013 [cited by applicant]
US 20150100957A1 · Botzer · 2015 [cited by applicant]
US 20150186172A1 · Thomas et al. · 2015 [cited by applicant]
US 20150237132A1 · Antony · 2015 [cited by applicant]
US 20160080451A1 · Morton et al. · 2016 [cited by applicant]
US 20170077964A1 · Guilford · 2017 [cited by examiner]
US 20170244593A1 · Rangasamy et al. · 2017 [cited by applicant]
US 20180060104A1 · Tarasuk-Levin et al. · 2018 [cited by applicant]
US 20180123895A1 · Khasnabish et al. · 2018 [cited by applicant]
US 20190041967A1 · Ananthakrishnan et al. · 2019 [cited by applicant]
US 20190065281A1 · Bernat · 2019 [cited by examiner]
US 20190243672A1 · Yadav et al. · 2019 [cited by applicant]
US 20200097280A1 · Simeonov et al. · 2020 [cited by applicant]
US 20210081216A1 · Komarov et al. · 2021 [cited by applicant]
US 20210326763A1 · Bernat · 2021 [cited by examiner]
US 20220385543A1 · Villasante Marcos et al. · 2022 [cited by applicant]
US 20220404798A1 · Amaro, Jr. et al. · 2022 [cited by applicant]
US 20230124947A1 · Mermoud et al. · 2023 [cited by applicant]
US 20240143372A1 · Currid et al. · 2024 [cited by applicant]
US 20240143408A1 · Currid · 2024 [cited by examiner]
US 20240179086A1 · Lee · 2024 [cited by examiner]
NCCL, “NVIDIA Collective Communication Library (NCCL) Documentation,” Retrieved from, <https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/index.html,> 2020, 9 Pages. [cited by applicant]
OPENUCX, “OPENUCX,” Retrieved from, <https://openucx.readthedocs.io/en/master/,> 2019, 2 Pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
IEEE, “IEEE Standard for 802.3,” IEEE Standard for Ethemetn, IEEE Computer Society, Dec. 28, 2012, 634 pages. [cited by applicant]
Wikipedia, “IEEE 802.11,” Wikipedia the Free Encyclopedia, https://en.wikipedia.org/wiki/IEEE_802.11, most recent edit Sep. 20, 2020 [retrieved Sep. 22, 2020], 15 pages. [cited by applicant]
Wikipedia, “IEEE 802.5,” Wikepedia The Free Encyclopedia, https://en.wikipedia.org/wiki/Token_Ring, Jan. 14, 2020, 12 pages. [cited by applicant]