IP Library › Granted Patent US 12,625,729
Granted Patent B2
US 12,625,729 · App. 17/752,790 · Granted May 12, 2026

Packet processing computations utilizing a pre-allocated memory function

Inventors: Janardhana Reddy Naredula (Hyderabad, IN); Naresh Kumar Bade (Hyderabad, IN)
Assignee: Microsoft Technology Licensing, LLC
G06F9/5016G06F9/5038G06F9/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,729
App. No.
17/752,790
Granted
May 12, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and computer-readable media for utilizing a new memory allocation function library called PmemMalloc. For example, the PmemMalloc library allocates pre-allocated, partitioned, and fixed shared memory blocks. In addition, by utilizing the PmemMalloc library, the memory allocation system described herein overcomes problems with persistence and enumeration that encumber existing malloc libraries. Indeed, the PmemMalloc library enables the memory allocation system to perform servicing computation in parallel across multiple CPU cores/threads, distribute computation equally among threads, prioritize servicing, among other improvements. Notably, the PmemMalloc library provides major constructs (e.g., persistence, enumeration, and debuggability) not available existing malloc libraries. Additionally, as detailed in this disclosure, the PmemMalloc library migrates various computations out of application-based packet processing to memory block-based deferred enumeration, which improves both packet processing and efficient use of CPU cores on a computing device.

Claims (46)

1 . A computer-implemented method for managing memory, comprising:

creating shared memory blocks that are pre-allocated, partitioned, and fixed in size, the shared memory blocks being accessible by multiple applications;

subsequent to completion of creating the shared memory blocks, receiving packet data at a network device that comprises an application that processes the packet data;

in response to receiving the packet data, processing the packet data by creating a memory heap within the shared memory blocks upon receiving the packet data, wherein the shared memory blocks are pre-allocated within a memory region, partitioned in the memory region, and fixed in size within the memory region, and wherein the shared memory blocks are accessible by multiple applications, wherein the memory heap persists across rebooting or restarting the application and does not involve operating integration;

generating allocation metadata for each block of the shared memory blocks, wherein the allocation metadata includes an object type of packet data portions stored in each block of the shared memory blocks;

attaching the allocation metadata to each corresponding block of the shared memory blocks; and

performing one or more memory operations to perform packet processing operations on the packet data and perform enumeration operations across the shared memory blocks based on the allocation metadata.

2 . The computer-implemented method of claim 1 , further comprising generating the allocation metadata for each block of the shared memory blocks before attaching the allocation metadata to corresponding blocks of the shared memory blocks.

3 . The computer-implemented method of claim 2 , wherein the allocation metadata for each block of the shared memory blocks comprises a timestamp, and a port number.

4 . The computer-implemented method of claim 3 , wherein the enumeration operations include zero blocking, servicing with data structure change, object expiry, port delete, and memory leak detection.

5 . The computer-implemented method of claim 4 , wherein performing the one or more memory operations across the shared memory blocks is further based on priority levels of the enumeration operations comprising a high-priority bulk operation, a low-priority operation, and an optional-priority operation.

6 . The computer-implemented method of claim 5 , wherein performing the one or more memory operations across the shared memory blocks further comprises processing high-priority bulk operations while deferring low-priority bulk operations until CPU idle time increases beyond a threshold level.

7 . The computer-implemented method of claim 4 , wherein performing the one or more memory operations across the shared memory blocks comprising:

filtering memory blocks from the shared memory blocks based on object types; and

performing an enumeration operation on the filtered memory blocks having a same object type.

8 . The computer-implemented method of claim 1 , wherein performing the one or more memory operations across the shared memory blocks comprising:

sharding the packet data across the shared memory blocks; and

performing both the packet processing operations and the enumeration operations across multiple CPU data threads utilizing the shared memory blocks.

9 . The computer-implemented method of claim 1 , wherein the memory heap of the shared memory blocks comprises memory that persists across rebooting the application.

10 . The computer-implemented method of claim 1 , wherein the memory heap of the shared memory blocks is pre-allocated without operating system integration.

11 . The computer-implemented method of claim 1 , further comprising receiving the packet data corresponding to the application from a network interface card before flows are converted into new objects.

12 . The computer-implemented method of claim 1 , wherein generating the memory heap of the shared memory blocks comprising pre-allocating and partitioning the shared memory blocks across a fixed number of data threads, a fixed number of CPUs, and a fixed amount of memory.

13 . The computer-implemented method of claim 1 , wherein CPU computations and the shared memory blocks occur on a data plane development server device.

14 . The computer-implemented method of claim 1 , wherein performing the one or more memory operations across the shared memory blocks based on the allocation metadata comprises filtering and selecting one or more shared memory blocks that include a same metadata object type for combined processing.

15 . A system comprising:

at least one processor; and

a non-transitory computer memory comprising instructions that, when executed by the at least one processor, cause the system to:

create shared memory blocks that are pre-allocated, partitioned, and fixed in size, the shared memory blocks being accessible by multiple applications;

subsequent to completion of creating the shared memory blocks, receive packet data at a network device that comprises an application that processes the packet data;

in response to receiving the packet data, process the packet data by creating a memory heap within the shared memory blocks upon receiving the packet data, wherein the shared memory blocks are pre-allocated within a memory region, partitioned in the memory region, and fixed in size within the memory region, and wherein the shared memory blocks are accessible by multiple applications, wherein the memory heap persists across rebooting or restarting the application and does not involve operating integration;

generate allocation metadata for each block of the shared memory blocks, wherein the allocation metadata includes an object type of packet data portions stored in each block of the shared memory blocks;

attach the allocation metadata to each corresponding block of the shared memory blocks; and

perform one or more memory operations to perform packet processing operations on the packet data and perform enumeration operations across the shared memory blocks based on the allocation metadata.

16 . The system of claim 15 , wherein performing the one or more memory operations across the shared memory blocks is further based on priority levels of the enumeration operations that comprise a high-priority bulk operation and a low-priority operation.

17 . The system of claim 15 , further comprising additional instructions that, when executed by the at least one processor, cause the system to receive the packet data corresponding to the application from a network interface card before flows are converted into new objects.

18 . The system of claim 15 , wherein generating the memory heap of the shared memory blocks comprising pre-allocating and partitioning the shared memory blocks across a fixed number of data threads, a fixed number of CPUs, and a fixed amount of memory.

19 . A non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause a computing device to:

create shared memory blocks that are pre-allocated, partitioned, and fixed in size, the shared memory blocks being accessible by multiple applications;

subsequent to completion of creating the shared memory blocks, receive packet data at a network device that comprises an application that processes the packet data;

in response to receiving the packet data, process the packet data by creating a memory heap within the shared memory blocks upon receiving the packet data, wherein the shared memory blocks are pre-allocated within a memory region, partitioned in the memory region, and fixed in size within the memory region, and wherein the shared memory blocks are accessible by multiple applications, wherein the memory heap persists across rebooting or restarting the application and does not involve operating integration;

generate allocation metadata for each block of the shared memory blocks, wherein the allocation metadata includes an object type of packet data portions stored in each block of the shared memory blocks;

attach the allocation metadata to each corresponding block of the shared memory blocks; and

perform one or more memory operations to perform packet processing operations on the packet data and perform enumeration operations across the shared memory blocks based on the allocation metadata.

20 . The non-transitory computer-readable storage medium of claim 19 , further comprising additional instructions that, when executed by the at least one processor, cause the computing device to generate the allocation metadata for each block of the shared memory blocks before attaching the allocation metadata to corresponding blocks of the shared memory blocks, and wherein:

the allocation metadata for each block of the shared memory blocks comprises a timestamp, and a port number; and

performing the one or more memory operations across the shared memory blocks is further based on priority levels of the enumeration operations comprising a high-priority bulk operation, a low-priority operation, and an optional-priority operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2022
From: BADE, NARESH KUMAR; NAREDULA, JANARDHANA REDDY
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061143/0648 →
Continuity (1)
Related Publication 20230385111A1 · Nov 30, 2023
References Cited (12)
US 10278112B1 · A · 2019 [cited by examiner]
US 20020095453A1 · Steensgaard · 2002 [cited by examiner]
US 20080209154A1 · Schneider · 2008 [cited by examiner]
US 20200177660A1 · Connor · 2020 [cited by examiner]
US 20200374744A1 · Liu · 2020 [cited by examiner]
US 20210011765A1 · Doshi · 2021 [cited by examiner]
US 20210089411A1 · Prasad · 2021 [cited by examiner]
US 20210271535A1 · Ritchie et al. · 2021 [cited by applicant]
US 20210342293A1 · Waddington · 2021 [cited by examiner]
US 20220114084A1 · Ritchie · 2022 [cited by examiner]
“DPDK Documentation”, Retrieved from: https://dpdk.readthedocs.io/_/downloads/en/v2.1.0/pdf/, Aug. 18, 2015, 528 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/016786”, Mailed Date: Jul. 10, 2023, 11 Pages. [cited by applicant]