IP Library Granted Patent US 10,325,343
Granted Patent B1
US 10,325,343 · App. 15/669,452 · Granted Jun 18, 2019

Topology aware grouping and provisioning of GPU resources in GPU-as-a-Service platform

Inventors: Junping Zhao (Beijing, CN); Zhi Ying (Shanghai, CN); Kenneth Durazzo (San Ramon, CA)
Assignee: EMC IP Holding Company LLC
G06T1/20H04L67/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,325,343
App. No.
15/669,452
Granted
Jun 18, 2019
Kind
B1
Abstract

Techniques are provided for implementing a graphics processing unit (GPU) service platform that is configured to provide topology aware grouping and provisioning of GPU resources for GPU-as-a-Service. A GPU server node receives a service request from a client system for GPU processing services provided by the GPU server node, wherein the GPU server node comprises a plurality of GPU devices. The GPU server node accesses a performance metrics data structure which comprises performance metrics associated with an interconnect topology of the GPU devices and hardware components of the GPU sever node. The GPU server node dynamically forms a group of GPU devices of the GPU server node based on the performance metrics of the accessed data structure, and provisions the dynamically formed group of GPU devices to the client system to handle the service request.

Claims (44)

1. A method, comprising:

receiving, by a graphics processing unit (GPU) server node, a service request from a client system for GPU processing services provided by the GPU server node, wherein the GPU server node comprises a plurality of GPU devices;

determining a hardware interconnect topology of the GPU server node, the hardware interconnect topology comprising information regarding determined interconnect paths between the GPU devices of the GPU server node and between the GPU devices and hardware components of the GPU server node;

accessing a performance metrics data structure which comprises performance metrics associated with a plurality of different interconnect path types that can be used to connect to GPU devices in a GPU server node topology, wherein the performance metrics comprise predefined priority scores that accord different priorities to the plurality of different interconnect path types;

dynamically forming a group of GPU devices for handling the service request received from the client system based at least in part on the determined hardware interconnect topology of the GPU server node and the performance metrics, wherein the group of GPU devices is dynamically formed at least in part by selecting one or more GPU devices of the GPU server node which are determined to be interconnected with higher priority interconnect paths as compared to other GPU devices of the GPU server node which are determined to be interconnected with lower priority interconnect paths; and

provisioning the dynamically formed group of GPU devices to the client system to handle the service request.

2. The method of claim 1 , further comprising accessing one or more quality of service policies associated with the client system, wherein the group of GPU devices is dynamically formed based at least in part on the one or more quality of service policies associated with the client system.

3. The method of claim 1 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect a GPU device and a network adapter which is utilized to connect to the GPU server node.

4. The method of claim 1 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect two GPU devices of the GPU server node.

5. The method of claim 1 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path which includes a processor socket-level link, wherein the priority accorded to the type of interconnect path which includes a processor socket-level link is lower than a priority accorded to a type of interconnect path which does not include a processor socket-level link.

6. The method of claim 1 , further comprising:

generating a system topology data structure comprising information regarding the determined hardware interconnect topology of the GPU server node; and

populating the system topology data structure with priority information for the determined interconnect paths included in the system topology data structure based on the predefined priority scores of the performance metrics data structure.

7. The method of claim 1 , wherein dynamically forming the group of GPU devices of the GPU server node comprises selecting one or more GPU devices that are already included as part of one or more other dynamically formed groups of GPU devices provisioned to other client systems.

8. An article of manufacture comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code is executable by a processor to implement a process comprising:

receiving, by a graphics processing unit (GPU) server node, a service request from a client system for GPU processing services provided by the GPU server node, wherein the GPU server node comprises a plurality of GPU devices;

determining a hardware interconnect topology of the GPU server node, the hardware interconnect topology comprising information regarding determined interconnect paths between the GPU devices of the GPU server node and between the GPU devices and hardware components of the GPU server node;

accessing a performance metrics data structure which comprises performance metrics associated with a plurality of different interconnect path types that can be used to connect to GPU devices in a GPU server node topology, wherein the performance metrics comprise predefined priority scores that accord different priorities to the plurality of different interconnect path types;

dynamically forming a group of GPU devices for handling the service request received from the client system based at least in part on the determined hardware interconnect topology of the GPU server node and the performance metrics, wherein the group of GPU devices is dynamically formed at least in part by selecting one or more GPU devices of the GPU server node which are determined to be interconnected with higher priority interconnect paths as compared to other GPU devices of the GPU server node which are determined to be interconnected with lower priority interconnect paths; and

provisioning the dynamically formed group of GPU devices to the client system to handle the service request.

9. The article of manufacture of claim 8 , further comprising program code that is executable by the processor to perform a method comprising accessing one or more quality of service policies associated with the client system, wherein the group of GPU devices is dynamically formed based at least in part on the one or more quality of service policies associated with the client system.

10. The article of manufacture of claim 8 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect a GPU device and a network adapter which is utilized to connect to the GPU server node.

11. The article of manufacture of claim 8 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect two GPU devices of the GPU server node.

12. The article of manufacture of claim 8 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path which includes a processor socket-level link, wherein the priority accorded to the type of interconnect path which includes a processor socket-level link is lower than a priority accorded to a type of interconnect path which does not include a processor socket-level link.

13. The article of manufacture of claim 8 , further comprising program code that is executable by the processor to perform a method comprising:

generating a system topology data structure comprising information regarding the determined hardware interconnect topology of the GPU server node; and

populating the system topology data structure with priority information for the determined interconnect paths included in the system topology data structure based on the predefined priority scores of the performance metrics data structure.

14. The article of manufacture of claim 8 , wherein dynamically forming the group of GPU devices of the GPU server node comprises selecting one or more GPU devices that are already included as part of one or more other dynamically formed groups of GPU devices provisioned to other client systems.

15. A graphics processing unit (GPU) server node, comprising:

a plurality of GPU devices;

a memory to store program instructions; and

a processor to execute the stored program instructions to cause the GPU server node to perform a process which comprises:

receiving a service request from a client system for GPU processing services provided by the GPU server node;

determining a hardware interconnect topology of the GPU server node, the hardware interconnect topology comprising information regarding determined interconnect paths between the GPU devices of the GPU server node and between the GPU devices and hardware components of the GPU server node;

accessing a performance metrics data structure which comprises performance metrics associated with a plurality of different interconnect path types that can be used to connect to GPU devices in a GPU server node topology, wherein the performance metrics comprise predefined priority scores that accord different priorities to the plurality of different interconnect path types;

dynamically forming a group of GPU devices for handling the service request received from the client system based at least in part on the determined hardware interconnect topology of the GPU server node and the performance metrics, wherein the group of GPU devices is dynamically formed at least in part by selecting one or more GPU devices of the GPU server node which are determined to be interconnected with higher priority interconnect paths as compared to other GPU devices of the GPU server node which are determined to be interconnected with lower priority interconnect paths; and

provisioning the dynamically formed group of GPU devices to the client system to handle the service request.

16. The GPU server node of claim 15 , therein the process performed by the GPU server node further comprises accessing one or more quality of service policies associated with the client system, wherein the group of GPU devices is dynamically formed based at least in part on the one or more quality of service policies associated with the client system.

17. The GPU server node of claim 15 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect a GPU device and a network adapter which is utilized to connect to the GPU server node.

18. The GPU server node of claim 15 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path that is used to connect two GPU devices of the GPU server node.

19. The GPU server node of claim 15 , wherein the performance metrics comprise a predefined priority score that is accorded to a type of interconnect path which includes a processor socket-level link, wherein the priority accorded to the type of interconnect path which includes a processor socket-level link is lower than a priority accorded to a type of interconnect path which does not include a processor socket-level link.

20. The GPU server node of claim 15 , therein the process performed by the GPU server node further comprises:

generating a system topology data structure comprising information regarding the determined hardware interconnect topology of the GPU server node; and

populating the system topology data structure with priority information for the determined interconnect paths included in the system topology data structure based on the predefined priority scores of the performance metrics data structure.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (043775/0082) Recorded May 20, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 060958/0468 →
RELEASE OF SECURITY INTEREST AT REEL 043772 FRAME 0750 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
Reel/Frame 058298/0606 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2018
From: ZHAO, JUNPING; YING, ZHI; DURAZZO, KENNETH
To: EMC IP HOLDING COMPANY LLC
Reel/Frame 046833/0828 →
PATENT SECURITY AGREEMENT (CREDIT) Recorded Sep 6, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 043772/0750 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Sep 6, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 043775/0082 →
Cited By (8)
US 12,229,581 US 12,266,030 US 12,399,745 US 12,471,238 US 12,608,227 US 12,632,314 US 12,694,602 US 12,705,061