IP Library Granted Patent US 12,591,510
Granted Patent B2
US 12,591,510 · App. 18/868,114 · Granted Mar 31, 2026

Systems and methods of allocating GPU memory

Inventor: Maarten Hoeben (San Jose, CA)
Assignee: ACTIVEVIDEO NETWORKS, LLC.
G06F12/023G06F9/5016G06F2212/1008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,510
App. No.
18/868,114
Granted
Mar 31, 2026
Kind
B2
Abstract

The server, initializes, for a third-party application executing on the server, an entirety of available GPU memory of a client device, including pre-allocating a plurality of blocks of GPU memory. During execution of the third-party application, the server receives a first request from the third-party application to store first data in the GPU memory of the client device, and, in response to the first request, frees a portion of a respective pre-allocated block of the plurality of pre-allocated blocks of GPU memory and stores the first data in the portion of the respective pre-allocated block. The server pre-allocates a new block of GPU memory of the client device, the new block comprising a complementary portion of the respective pre-allocated block such that, after pre-allocating the new block of GPU memory, the entirety of available GPU memory of the client device remains allocated.

Claims (41)

1 . A method, comprising:

at a server system hosting a virtual machine executing a third-party application, the server system in communication with a physical client device:

initializing, for the third-party application, an entirety of available GPU memory of the client device, including pre-allocating a plurality of blocks of GPU memory;

during execution of the third-party application:

receiving a first request from the third-party application to store first data in the GPU memory of the client device;

in response to the first request:

freeing a portion of a respective pre-allocated block of the plurality of pre-allocated blocks of GPU memory;

storing the first data in the portion of the respective pre-allocated block; and

pre-allocating a new block of GPU memory of the client device, the new block comprising a complementary portion of the respective pre-allocated block such that, after pre-allocating the new block of GPU memory, the entirety of available GPU memory of the client device remains allocated.

2 . The method of claim 1 , further comprising:

storing a map of pre-allocated blocks, the map including an identifier and size of each of the plurality of pre-allocated blocks; and

in response to the first request from the third-party application to store the first data in the GPU memory of the client device, updating the map to include the pre-allocated new block of memory.

3 . The method of claim 1 , wherein the pre-allocated blocks have a maximum size.

4 . The method of claim 1 , wherein the pre-allocated blocks do not include data for the third-party application.

5 . The method of any of claims 1-4 , further including determining a position, within the respective pre-allocated block, of the portion of the respective pre-allocated block in which the first data is stored using a known management scheme of the physical client device.

6 . The method of any of claims 1-4 , wherein pre-allocating the plurality of blocks of GPU memory comprises iteratively pre-allocating blocks of decreasing size until the entirety of the GPU memory is pre-allocated.

7 . The method of any of claims 1-4 , further comprising:

receiving a second request from the third-party application to store second data in the GPU memory of the client device;

in response to the second request:

determining that the second data is larger than any currently pre-allocated blocks of GPU memory;

in accordance with the determination that the second data is larger than any currently pre-allocated blocks of GPU memory, moving the first data to a different pre-allocated block; and

storing the second data in GPU memory freed by moving the first data to the different pre-allocated block.

8 . The method of any of claims 1-4 , wherein the physical client device does not include a memory manager for the GPU memory.

9 . A non-transitory computer-readable storage medium storing one or more programs for execution by a server system executing a third-party application, the server system in communication with a client device, the one or more programs including instructions for:

initializing, for the third-party application, an entirety of available GPU memory of the client device, including pre-allocating a plurality of blocks of GPU memory;

during execution of the third-party application:

receiving a first request from the third-party application to store first data in the GPU memory of the client device;

in response to the first request:

freeing a portion of a respective pre-allocated block of the plurality of pre-allocated blocks of GPU memory;

storing the first data in the portion of the respective pre-allocated block; and

pre-allocating a new block of GPU memory of the client device, the new block comprising a complementary portion of the respective pre-allocated block such that, after pre-allocating the new block of GPU memory, the entirety of available GPU memory of the client device remains allocated.

10 . A server system hosting a virtual machine executing a third-party application, the server system in communication with a physical client device, comprising:

one or more processors; and

memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for:

initializing, for the third-party application, an entirety of available GPU memory of the client device, including pre-allocating a plurality of blocks of GPU memory;

during execution of the third-party application:

receiving a first request from the third-party application to store first data in the GPU memory of the client device;

in response to the first request:

freeing a portion of a respective pre-allocated block of the plurality of pre-allocated blocks of GPU memory;

storing the first data in the portion of the respective pre-allocated block; and

pre-allocating a new block of GPU memory of the client device, the new block comprising a complementary portion of the respective pre-allocated block such that, after pre-allocating the new block of GPU memory, the entirety of available GPU memory of the client device remains allocated.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2025
From: HOEBEN, MAARTEN
To: ACTIVEVIDEO NETWORKS, LLC
Reel/Frame 070066/0748 →
Continuity (2)
Provisional Application 63346780 · May 27, 2022
Related Publication 20250328459A1 · Oct 23, 2025
References Cited (11)
US 8984167B1 · Diard · 2015 [cited by examiner]
US 10846096B1 · Chung et al. · 2020 [cited by applicant]
US 20090031289A1 · Michael · 2009 [cited by applicant]
US 20090313451A1 · Inoue et al. · 2009 [cited by applicant]
US 20150188992A1 · Ayanam et al. · 2015 [cited by applicant]
US 20170034297A1 · Waheed · 2017 [cited by applicant]
US 20200098082A1 · Gutierrez et al. · 2020 [cited by applicant]
US 20200134208A1 · Pappachan et al. · 2020 [cited by applicant]
US 20210011773A1 · Garg · 2021 [cited by examiner]
Lin et al. “A New Non-Blocking Approach on GPU Dynamical Memory Management.” In: International Workshop on Computational Science and Engineering, Oct. 14-17, 2013, [online] [retrived on Jul. 26, 2023] Retrieved from the… [cited by applicant]
International Search Report in International Application No. PCT/US23/23193, mailed Aug. 22, 2023. [cited by applicant]