IP Library Granted Patent US 10,728,091
Granted Patent B2
US 10,728,091 · App. 15/945,412 · Granted Jul 28, 2020

Topology-aware provisioning of hardware accelerator resources in a distributed environment

Inventors: Junping Zhao (Beijing, CN); Yunfan Han (Austin, TX)
Assignee: EMC IP Holding Company LLC
H04L41/0806G06F9/5011G06N3/08H04L41/12H04L41/14H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,728,091
App. No.
15/945,412
Granted
Jul 28, 2020
Kind
B2
Abstract

Techniques are provided for topology-aware provisioning of computing resources in a distributed heterogeneous environment. For example, a method includes: receiving a service request from a client system to perform a data processing job in a server cluster; determining candidate accelerator devices that reside in server nodes of the server cluster, which can be utilized to perform the data processing job; determining a connection topology of each candidate accelerator device within the server nodes, and a performance ranking of each connection topology; utilizing the determined performance ranking of each connection topology to select a group of accelerator devices among the candidate accelerator devices, which can be provisioned to perform the data processing job, wherein the selected group of accelerator devices include candidate accelerator devices with connection topologies that have matching performance rankings; and scheduling and provisioning the selected group of accelerator devices to execute the data processing job.

Claims (43)

1. A method, comprising:

receiving, by a control server node, a service request from a client system to perform a data processing job in a server cluster managed by the control server node;

determining, by the control server node, a set of candidate accelerator devices that resides in one or more server nodes of the server cluster, which can be utilized to perform the data processing job;

determining, by the control server node, connection topologies between respective pairs of accelerator devices within the set of candidate accelerator devices, and a performance ranking of each connection topology, wherein a given connection topology between a given pair of accelerator devices comprises information regarding a type of interconnect path between the given pair of accelerator devices, and wherein the performance ranking of the given connection topology comprises a rank score that is accorded to the given connection topology among a plurality of different rank scores that are accorded to respective different connection topologies;

utilizing, by the control server node, the determined performance ranking of each connection topology to select a group of accelerator devices among the set of candidate accelerator devices, which can be provisioned to perform the data processing job, wherein the selected group of accelerator devices includes one or more pairs of accelerator devices with connection topologies that have matching performance rankings; and

configuring, by the control server node, the selected group of accelerator devices having the connection topologies with matching performance rankings to execute the data processing job.

2. The method of claim 1 , wherein the service request comprises one or more resource demands, and wherein determining the set of candidate accelerator devices comprises determining the set of candidate accelerator devices which satisfy the one or more resource demands.

3. The method of claim 2 , wherein the one or more resource demands comprise one of (i) a requested number of accelerator devices to be provisioned for the data processing job and (ii) a type of accelerator device to be provisioned for the data processing job.

4. The method of claim 1 , wherein the selected group of accelerator devices includes candidate accelerator devices with a same connection topology.

5. The method of claim 1 , wherein the selected group of accelerator devices includes candidate accelerator devices with a same highest performance ranking.

6. The method of claim 1 , wherein the selected group of accelerator devices includes candidate accelerator devices with different connection topologies, wherein the different connection topologies have similar performance rankings which meet a predefined similarity matching rule.

7. The method of claim 1 , further comprising determining, by the control server node, a communication sequence of the selected group of accelerator devices for configuring the selected group of accelerator devices in a logical communication ring.

8. The method of claim 7 , wherein the selected group of accelerator devices is configured in the logical communication ring and provisioned to perform a distributed deep learning model training job using a Ring AllReduce protocol.

9. The method of claim 7 , wherein determining the communication sequence comprises:

determining a current bandwidth usage of communication buses to which the selected group of accelerator devices is connected; and

determining the communication sequence of the selected group of accelerator devices based on the determined bandwidth usage to optimize usage of the communication buses.

10. An article of manufacture comprising a processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code is executable by a processor to implement a process comprising:

receiving, by a control server node, a service request from a client system to perform a data processing job in a server cluster managed by the control server node;

determining, by the control server node, a set of candidate accelerator devices that resides in one or more server nodes of the server cluster, which can be utilized to perform the data processing job;

determining, by the control server node, connection topologies between respective pairs of accelerator devices within the set of candidate accelerator devices, and a performance ranking of each connection topology, wherein a given connection topology between a given pair of accelerator devices comprises information regarding a type of interconnect path between the given pair of accelerator devices, and wherein the performance ranking of the given connection topology comprises a rank score that is accorded to the given connection topology among a plurality of different rank scores that are accorded to respective different connection topologies;

utilizing, by the control server node, the determined performance ranking of each connection topology to select a group of accelerator devices among the set of candidate accelerator devices, which can be provisioned to perform the data processing job, wherein the selected group of accelerator devices includes one or more pairs of accelerator devices with connection topologies that have matching performance rankings; and

configuring, by the control server node, the selected group of accelerator devices having the connection topologies with matching performance rankings to execute the data processing job.

11. The article of manufacture of claim 10 , wherein the service request comprises one or more resource demands, wherein the one or more resource demands comprise one of (i) a requested number of accelerator devices to be provisioned for the data processing job and (ii) a type of accelerator device to be provisioned for the data processing job, and wherein determining the set of candidate accelerator devices comprises determining the set of candidate accelerator devices which satisfy the one or more resource demands.

12. The article of manufacture of claim 10 , wherein the selected group of accelerator devices includes candidate accelerator devices with a same connection topology.

13. The article of manufacture of claim 10 , wherein the selected group of accelerator devices includes candidate accelerator devices with a same highest performance ranking.

14. The article of manufacture of claim 10 , wherein the selected group of accelerator devices includes candidate accelerator devices with different connection topologies, wherein the different connection topologies have similar performance rankings which meet a predetermined similarity matching rule.

15. The article of manufacture of claim 10 , further comprising executable program code for determining, by the control server node, a communication sequence of the selected group of accelerator devices for configuring the selected group of accelerator devices in a logical communication ring.

16. The article of manufacture of claim 15 , wherein the selected group of accelerator devices is configured in the logical communication ring and provisioned to perform a distributed deep learning model training job using a Ring AllReduce protocol.

17. The article of manufacture of claim 15 , wherein determining the communication sequence comprises:

determining a current bandwidth usage of communication buses to which the selected group of accelerator devices is connected; and

determining the communication sequence of the selected group of accelerator devices based on the determined bandwidth usage to optimize usage of the communication buses.

18. A system, comprising:

a server cluster comprising a plurality of server nodes, wherein the server nodes comprise accelerator devices;

a control server node comprising a memory to store program instructions, and a processor to execute the stored program instructions to cause the control server node to perform a process which comprises:

receiving a service request from a client system to perform a data processing job in a server cluster managed by the control server node;

determining a set of candidate accelerator devices that resides in one or more server nodes of the server cluster, which can be utilized to perform the data processing job;

determining connection topologies between respective pairs of accelerator devices within the set of candidate accelerator devices, and a performance ranking of each connection topology, wherein a given connection topology between a given pair of accelerator devices comprises information regarding a type of interconnect path between the given pair of accelerator devices, and wherein the performance ranking of the given connection topology comprises a rank score that is accorded to the given connection topology among a plurality of different rank scores that are accorded to respective different connection topologies;

utilizing the determined performance ranking of each connection topology to select a group of accelerator devices among the set of candidate accelerator devices, which can be provisioned to perform the data processing job, wherein the selected group of accelerator devices includes one or more pairs of accelerator devices with connection topologies that have matching performance rankings; and

configuring the selected group of accelerator devices having the connection topologies with matching performance rankings to execute the data processing job.

19. The system of claim 18 , wherein the selected group of accelerator devices includes candidate accelerator devices with one of (i) a same highest performance ranking, and (ii) different connection topologies, wherein the different connection topologies have similar performance rankings which meet a predetermined similarity matching rule.

20. The system of claim 18 , wherein the process performed by the control server node further comprises:

determining a communication sequence of the selected group of accelerator devices for configuring the selected group of accelerator devices in a logical communication ring, wherein the selected group of accelerator devices is configured in the logical communication ring and provisioned to perform a distributed deep learning model training job using a Ring AllReduce protocol;

wherein determining the communication sequence comprises determining a current bandwidth usage of communication buses to which the selected group of accelerator devices is connected, and determining the communication sequence of the selected group of accelerator devices based on the determined bandwidth usage to optimize usage of the communication buses.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (046366/0014) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 060450/0306 →
RELEASE OF SECURITY INTEREST AT REEL 046286 FRAME 0653 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 058298/0093 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Jun 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 046366/0014 →
PATENT SECURITY AGREEMENT (CREDIT) Recorded Jun 1, 2018
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 046286/0653 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2018
From: ZHAO, JUNPING; HAN, YUNFAN
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 045438/0001 →
Cited By (4)
US 12,236,248 US 12,255,951 US 12,375,554 US 12,638,903