IP Library Patent Application 17536018
Patent Application
App. No. 17/536,018

SCHEDULING SYSTEM FOR COMPUTATIONAL WORK ON HETEROGENEOUS HARDWARE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/536,018
Abstract

The technology includes methods, processes, and systems for virtualizing graphics processing unit (GPU) memory. Example embodiments of the technology include managing an amount of GPU memory used by one or more processes, such as Application Programming Interfaces (APIs), that directly or indirectly impact one or more other processes running on the same GPU. Managing and/or virtualizing the amount of GPU memory may ensure that an end user does not receive a GPU out-of-memory error because the API request is impacted by the processing of other API requests. A virtual machine with access to a GPU may be organized with one or more job slots that are configured to specify the number of processes that are able to run concurrently on a specific virtual machine. A process may be configured on each virtual machine running a software program or API and is used to schedule work based on GPU memory requirements.

Claims (57)

1 . A computer-implemented method, comprising:

under the control of one or more computer systems configured with executable instructions,

managing an amount of Graphics Processing Unit (GPU) memory used by one or more processes, wherein the one or more processes directly or indirectly impact one or more other processes running on the GPU;

organizing a host machine with access to the GPU according to one or more request slots configured to specify a number of processes that are available to be processed by the GPU; and

scheduling the one or more processes based at least in part on the GPU memory.

2 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes Application Programming Interface (API) requests, complex-interaction calculations, neural networks, artificial intelligence, or other computation-intensive application.

3 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes:

optimizing a queue, in response to receiving two or more requests at or around a same time;

determining an amount of time to process a first request of the two or more requests;

determining an amount of time to process a second request of the two or more requests; and

ordering the first request and the second request in a queue, wherein the queue is configured to store the two or more requests.

4 . The computer-implemented method of claim 1 , wherein the computer-implemented method further includes:

receiving, from a client device, a request to add a persistent slot;

scheduling one process of the one or more processes in the persistent slot; and

determining if the one process executes one or more child processes.

5 . A system, comprising:

at least one computing device configured to implement one or more services, wherein the one or more services are configured to:

manage an amount of Graphics Processing Unit (GPU) memory used by one or more processes, wherein the one or more processes directly or indirectly impact one or more other processes running on the GPU;

organize a host machine with access to the GPU according to one or more request slots configured to specify a number of processes that are available to be processed by the GPU; and

to schedule the one or more processes based at least in part on the GPU memory.

6 . The system of claim 5 , wherein the one or more processes are received from one or more client devices.

7 . The system of claim 5 , wherein the at least one computing device is further configured to:

receive, from a client device, a request to add a persistent slot;

schedule one process of the one or more processes in the persistent slot; and

determine if the one process executes one or more child processes.

8 . The system of claim 5 , wherein the at least one computing device is further configured to:

optimize a queue, in response to receiving two or more requests at or around a same time;

determine an amount of time to process a first request of the two or more requests;

determine an amount of time to process a second request of the two or more requests; and

order the first request and the second request in a queue, wherein the queue is configured to store the two or more requests.

9 . The system of claim 8 , wherein the at least one computing device is further configured to:

receive at least one of information, input, and data associated with the one or more processes; and

store the at least one of information, input, and data in a database operably connected to the queue.

10 . The system of claim 5 , wherein the host machine is a virtual machine operably connected to a GPU.

11 . The system of claim 5 , wherein one or more processes include Application Programming Interface (API) requests, complex-interaction calculations, neural networks, artificial intelligence, or other computation-intensive applications.

12 . The system of claim 5 , wherein the at least one computing device is further configured to:

launch a new host machine;

create a new slot on an existing host machine; and

designate additional resources to the host machine from a pool of host machines operably connected to the GPU.

13 . A non-transitory computer-readable storage medium having stored thereon executable instructions that, when executed by one or more processors of a computer system, cause the computer system to at least:

receive, from a client device, a request to process data associated with the request;

schedule the request to one or more resources of a Graphics Processing Unit (GPU);

identify an amount of GPU resources being available to process the request using at least the data associated with the request; and

assign the request to the GPU.

14 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to provide, to the client device, the processed request.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the instructions that cause the computer system to provide the processed request further include instructions that cause the computer system to maintain the request assigned to the GPU after the processed request is provided to the client device.

16 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to maintain at least one of the request, the data associated with the request, and the processed request.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to:

receive one or more additional requests, the one or more additional requests including new data associated with the request;

retrieve the request maintained by the computer system; and

process the one or more additional requests based on the request.

18 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions that cause the computer system to schedule the request further include instructions that, when executed by the one or more processors, cause the computer system to schedule the request in a slot of the GPU, wherein the slot of the GPU is initialized based on the amount of GPU resources being available.

19 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions that cause the computer system to identify the amount of GPU resources being available to process the request further include instructions that, when executed by the one or more processors, cause the computer system to determine a status of one or more slots of the GPU, wherein the one or more slots of the GPU are configured to receive at least the request and the data associated with the request in order to process the request.

20 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions further comprise instructions that, when executed by the one or more processors, cause the computer system to:

determine if one or more other requests are assigned to the GPU;

identify the one or more other requests assigned to the GPU to be evicted from the GPU; and

evict the one or more other requests from the GPU.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Apr 7, 2025
From: CITIBANK, N.A.
To: DATAROBOT, INC.; ALGORITHMIA, INC.; DULLES RESEARCH, LLC
Reel/Frame 070750/0866 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2024
From: ALGORITHMIA, INC.
To: DATAROBOT, INC.
Reel/Frame 067633/0351 →
SECURITY INTEREST Recorded Mar 22, 2023
From: DATAROBOT, INC.; ALGORITHMIA, INC.; DULLES RESEARCH, LLC
To: CITIBANK, N.A.
Reel/Frame 063263/0926 →