IP Library › Granted Patent US 9,830,677
Granted Patent B2
US 9,830,677 · App. 15/059,580 · Granted Nov 28, 2017

Graphics processing unit resource sharing

Inventors: Anshul Gandhi (White Plains, NY); Hui Lei (Scarsdale, NY); Jayaram Kallapalayam Radhakrishnan (Mount Kisco, NY); Charles O. Schulz (Ridgefield, CT); Shu Tao (Irvington, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06T1/20G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,830,677
App. No.
15/059,580
Granted
Nov 28, 2017
Kind
B2
Abstract

Examples of GPU resource sharing among applications are disclosed. In one example, a method includes receiving a first request from a first application of the plurality of applications for first requested GPU resources, and receiving a second request from a second application of the plurality of applications for second GPU resources. The method also includes, responsive to determining that the first requested GPU resources are available, allocating a first slice of the GPU resources with a first requested amount of resources to the first application and, responsive to determining that the second requested GPU resources are available, allocating a second slice of the GPU resources with a second requested amount of resources to the second application. Further, the method includes enabling the first application and the second application to execute concurrently within the first slice of the GPU and the second slice of the GPU respectively.

Claims (45)

1. A computer-implemented method for sharing resources of a graphics processing unit (GPU) among a plurality of applications, the method comprising:

receiving a first request from a first application of the plurality of applications for first requested GPU resources, the GPU resources comprising a processor and a memory;

receiving a second request from a second application of the plurality of applications for second GPU resources;

responsive to determining that the first requested GPU resources are available, allocating a first slice of the GPU resources with a first requested amount of resources to the first application;

responsive to determining that the second requested GPU resources are available, allocating a second slice of the GPU resources with a second requested amount of resources to the second application; and

enabling the first application and the second application to execute concurrently within the first slice of the GPU and the second slice of the GPU respectively,

wherein allocating the first application to the first slice and allocating the second application to the second slice is performed according to a fairness policy, wherein the fairness policy establishes that an allocated percent of the memory is less than or equal to an allocated percent of the processor.

2. The computer-implemented method of claim 1 , wherein the first slice is inaccessible to the second application and wherein the second slice is inaccessible to the first application.

3. The computer-implemented method of claim 1 , further comprising:

receiving a third request from a third application of the plurality of applications for third requested GPU resources; and

responsive to determining that the third requested GPU resources are available, allocating a third slice of the GPU resources with a third requested amount of resources to the third application.

4. The computer-implemented method of claim 3 , further comprising:

enabling the third application to execute on the third slice concurrently with the first application executing on the first slice and the second application executing on the second slice.

5. The computer-implemented method of claim 1 , wherein receiving the first request, receiving the second request, allocating the first slice to a first container, and allocating the second slice to a second container are performed by a gatekeeper.

6. The computer-implemented method of claim 1 , wherein the processor comprises multiple GPU cores, wherein each GPU core of the multiple GPU cores comprises a plurality of hardware threads.

7. The computer-implemented method of claim 1 , wherein the first request is a request for at least one of a desired minimum and a desired maximum amount of resources, and wherein the second request is a request for at least one of a desired minimum and a desired maximum amount of resources.

8. A system for sharing resources of a graphics processing unit (GPU) among a plurality of applications, the system comprising:

a processor in communication with one or more types of memory, the processor configured to:

receive a first request from a first application of the plurality of applications for first requested GPU resources, the GPU resources comprising a processor and a memory;

receive a second request from a second application of the plurality of applications for second GPU resources;

responsive to determining that the first requested GPU resources are available, allocate a first slice of the GPU resources with a first requested amount of resources to the first application;

responsive to determining that the second requested GPU resources are available, allocate a second slice of the GPU resources with a second requested amount of resources to the second application; and

enable the first application and the second application to execute concurrently within the first slice of the GPU and the second slice of the GPU respectively,

wherein allocating the first application to the first slice and allocating the second application to the second slice is performed according to a fairness policy, wherein the fairness policy establishes that an allocated percent of the memory is less than or equal to an allocated percent of the processor.

9. The system of claim 8 , wherein the first slice is inaccessible to the second application and wherein the second slice is inaccessible to the first application.

10. The system of claim 8 , wherein the processor is further configured to:

receive a third request from a third application of the plurality of applications for third requested GPU resources; and

responsive to determining that the third requested GPU resources are available, allocate a third slice of the GPU resources with a third requested amount of resources to the third application.

11. The system of claim 10 , wherein the processor is further configured to:

enable the third application to execute on the third slice concurrently with the first application executing on the first slice and the second application executing on the second slice.

12. The system of claim 8 , wherein receiving the first request, receiving the second request, allocating the first slice to a first container, and allocating the second slice to a second container are performed by a gatekeeper.

13. The system of claim 8 , wherein the processor comprises multiple GPU cores, wherein each GPU core of the multiple GPU cores comprises a plurality of hardware threads.

14. The system of claim 8 , wherein the first request is a request for at least one of a desired minimum and a desired maximum amount of resources, and wherein the second request is a request for at least one of a desired minimum and a desired maximum amount of resources.

15. A computer program product for sharing resources of a graphics processing unit (GPU) among a plurality of applications, the computer program product comprising:

a non-transitory storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising:

receiving a first request from a first application of the plurality of applications for first requested GPU resources, the GPU resources comprising a processor and a memory;

receiving a second request from a second application of the plurality of applications for second GPU resources;

responsive to determining that the first requested GPU resources are available, allocating a first slice of the GPU resources with a first requested amount of resources to the first application;

responsive to determining that the second requested GPU resources are available, allocating a second slice of the GPU resources with a second requested amount of resources to the second application; and

enabling the first application and the second application to execute concurrently within the first slice of the GPU and the second slice of the GPU respectively,

wherein allocating the first application to the first slice and allocating the second application to the second slice is performed according to a fairness policy, wherein the fairness policy establishes that an allocated percent of the memory is less than or equal to an allocated percent of the processor.

16. The computer program product of claim 15 , wherein the first slice is inaccessible to the second application and wherein the second slice is inaccessible to the first application.

17. The computer program product of claim 15 , wherein the method further comprises:

receiving a third request from a third application of the plurality of applications for third requested GPU resources; and

responsive to determining that the third requested GPU resources are available, allocating a third slice of the GPU resources with a third requested amount of resources to the third application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2016
From: GANDHI, ANSHUL; LEI, HUI; RADHAKRISHNAN, JAYARAM KALLAPALAYAM; SCHULZ, CHARLES O.; TAO, SHU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 037883/0378 →
Continuity (1)
Related Publication 20170256017A1 · Sep 7, 2017