IP Library › Granted Patent US 12,386,648
Granted Patent B2
US 12,386,648 · App. 17/836,927 · Granted Aug 12, 2025

Resource allocation in virtualized environments

Inventors: Marjan Radi (San Jose, CA); Dejan Vucinic (San Jose, CA)
Assignee: Western Digital Technologies, Inc.
G06F9/45558G06F9/5077G06F2009/45583G06F2009/45595
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,648
App. No.
17/836,927
Granted
Aug 12, 2025
Kind
B2
Abstract

A node includes a shared memory for a distributed memory system. A Virtual Switch (VS) controller establishes different flows of packets between at least one Virtual Machine (VM) running at the node and one or more other VMs running at the node or at another node. Requests to access the shared memory are queued in submission queues in a kernel space and processed requests are queued in completion queues in the kernel space. Indications of queue occupancy are determined for at least one queue and one or more memory request rates are set for at least one application based at least in part on the determined indications of queue occupancy. In another aspect, flow metadata is generated for each flow and at least one of the set one or more respective memory request rates and one or more respective resource allocations is adjusted for the at least one application.

Claims (75)

1. A node, comprising:

at least one memory configured to be used at least in part as a shared memory shared by a plurality of user space applications in a distributed memory system in a network;

a network interface configured to communicate with one or more other nodes in the network; and

at least one processor configured, individually or in combination, to:

execute a Virtual Switching (VS) controller in a user space of the at least one memory, the VS controller configured to route a flow of packets between a Virtual Machine (VM) on the node and a different VM on the node or on a different node of the one or more other nodes, wherein the flow of packets is initiated by a user space application executed by the VM and includes at least one memory request to access the shared memory; and

execute at least one module in a kernel space of the at least one memory, the at least one module configured, individually or in combination, to:

queue memory requests to access the shared memory in a plurality of submission queues in the kernel space, wherein each submission queue of the plurality of submission queues corresponds to a different user space application initiating requests to access data in the shared memory;

queue processed memory requests to access the shared memory in one or more completion queues in the kernel space; and

provide one or more respective indications of queue occupancy for at least one queue of the plurality of submission queues and the one or more completion queues; and

wherein during the execution of the user space application, the user space application transmits the at least one memory request according to a memory request rate that is based at least in part on the one or more respective indications of queue occupancy for the at least one queue.

2. The node of claim 1 , wherein the at least one processor is further configured, individually or in combination, to execute a VS kernel module in the kernel space to provide the one or more respective indications of queue occupancy.

3. The node of claim 1 , wherein the at least one processor is further configured, individually or in combination, to execute a VS kernel module in the kernel space, the VS kernel module configured to, based on the memory request rate, add at least one of a congestion indicator and the memory request rate to one or more packets of the flow of packets.

4. The node of claim 1 , wherein the at least one processor is further configured, individually or in combination, to generate flow metadata for the flow of packets representing at least one of:

a read request frequency of memory requests to read data from the shared memory,

a write request frequency of memory requests to write data in the shared memory,

a read to write ratio of reading data from the shared memory to writing data in the shared memory,

a usage of an amount of the shared memory,

a usage of the at least one processor,

a packet reception rate, and

at least one indication of queue occupancy for a corresponding submission queue and a corresponding completion queue for the flow of packets.

5. The node of claim 4 , wherein the at least one processor is further configured, individually or in combination, to add at least part of the flow metadata to at least one packet sent from the node to at least one of a node of the one or more other nodes and a network controller in the network,

wherein a resource usage of a plurality of nodes in the network is based at least in part on the flow metadata added to the at least one packet.

6. The node of claim 1 , wherein the VS controller is further configured to:

receive flow metadata from a VS kernel module in the kernel space representing usage of at least one resource of the node by different flows of packets initiated by at least one user space application, wherein

at least one of one or more respective memory request rates and one or more respective resource allocations for the at least one user space application is based at least in part on the received flow metadata.

7. The node of claim 1 , wherein the VS controller is further configured to:

store old flow metadata received from a VS kernel module in the kernel space representing previous usage of at least one resource of the node by one or more previous flows of packets, wherein the memory request rate and a resource allocation for the user space application is based at least in part on an estimated future resource utilization of the user space application, and wherein the estimated future resource utilization is

based at least in part on the stored old flow metadata.

8. The node of claim 1 , wherein the plurality of submission queues includes Non-Volatile Memory express (NVMe) submission queues to access the shared memory.

9. The node of claim 1 , wherein a processor of the network interface is configured to:

provide the one or more respective indications of queue occupancy for the at least one queue of the plurality of submission queues and the one or more completion queues; and

send the one or more respective indications of queue occupancy to a different processor of the at least one processor that is configured to execute the VS controller.

10. A method, comprising:

executing a Virtual Switching (VS) controller in a user space of at least one memory of a node having at least one processor, the VS controller configured to establish different flows of packets between at least one Virtual Machine (VM) on the node and one or more other VMs on the node or on one or more other nodes in communication with the node via a network, wherein the different flows of packets include a plurality of memory requests of at least one user space application executing on at least one of the at least one VM and the one or more other VMs;

queuing memory requests from user space applications to access a shared memory of the at least one memory of the node in a plurality of submission queues in a kernel space of the at least one memory, wherein each submission queue of the plurality of submission queues corresponds to a different user space application initiating respective memory requests to access data in the shared memory;

queuing processed memory requests to access the shared memory in one or more completion queues in the kernel space;

based at least in part on at least one of the memory requests queued in the plurality of submission queues and the processed memory requests queued in the one or more completion queues, generating flow metadata for each of the different flows of packets, the flow metadata representing usage of at least one resource of the node by each of the different flows of packets; and

executing a user space application of the at least one user space application based at least in part on the generated flow metadata, wherein at least one of a memory request rate and one or more resource allocations for the user space application is based at least in part on the generated flow metadata.

11. The method of claim 10 , further comprising:

providing one or more respective indications of queue occupancy for at least one queue of the plurality of submission queues and the one or more completion queues; and

wherein the flow metadata is based at least in part on the one or more respective indications of queue occupancy for the at least one queue.

12. The method of claim 10 , wherein the flow metadata for each of the different flows of packets represents at least one of:

a read request frequency of memory requests to read data from the shared memory,

a write request frequency of memory requests to write data in the shared memory,

a read to write ratio of reading data from the shared memory to writing data in the shared memory,

a usage of an amount of the shared memory,

a usage of the at least one processor,

a packet reception rate,

an indication of the number of memory requests queued in at least one submission queue of the plurality of submission queues, and

at least one indication of queue occupancy for a corresponding submission queue and a corresponding completion queue for the flow of packets.

13. The method of claim 10 , further comprising:

executing a VS kernel module in the kernel space of the at least one memory; and

using the VS kernel module to generate the flow metadata for the different flows of packets.

14. The method of claim 10 , further comprising:

adding at least part of the flow metadata to at least one packet sent from the node to at least one of a node of the one or more other nodes in the network and a network controller in the network,

wherein a resource usage of a plurality of nodes in the network is based at least in part on the flow metadata added to the at least one packet.

15. The method of claim 10 , further comprising:

storing old flow metadata received from a VS kernel module in the kernel space representing previous usage of at least one resource of the node by one or more previous flows of packets, wherein at least one of the memory request rate and a resource allocation for the user space application is based at least in part on an estimated future resource utilization of the user space application, and wherein the estimated future resource utilization is

based at least in part on the stored old flow metadata.

16. The method of claim 10 , wherein the plurality of submission queues includes Non-Volatile Memory express (NVMe) submission queues to access the shared memory.

17. The method of claim 10 , further comprising:

using a processor of a network interface of the node to provide one or more respective indications of queue occupancy for at least one queue of the plurality of submission queues and the one or more completion queues; and

sending the one or more respective indications of queue occupancy to a different processor of the at least one processor that is configured to execute the VS controller.

18. A node, comprising:

at least one memory configured to be used at least in part as a shared memory shared by a plurality of user space applications in a distributed memory system in a network;

a network interface configured to communicate with one or more other nodes in the network;

a user space means for executing a Virtual Switch (VS) controller in a user space of the at least one memory, the VS controller configured to route a flow of packets between a Virtual Machine (VM) on the node and a different VM on the node or on a different node of the one or more other nodes, wherein the flow of packets is initiated by a user space application executed by the VM and includes at least one memory request to access the shared memory; and

a kernel space means for:

queuing memory requests to access the shared memory in a plurality of submission queues in a kernel space of the at least one memory, wherein each submission queue of the plurality of submission queues corresponds to a different user space application initiating requests to access data in the shared memory;

queuing processed memory requests to access the shared memory in one or more completion queues in the kernel space; and

providing one or more respective indications of queue occupancy for at least one queue of the plurality of submission queues and the one or more completion queues; and

wherein during the execution of the user space application, the user space application transmits the at least one memory request according to a memory request rate that is based at least in part on the one or more respective indications of queue occupancy for the at least one queue.

19. The node of claim 18 , wherein the kernel space means is further for, based at least in part on at least one of the memory requests queued in the plurality of submission queues and the processed memory requests queued in the one or more completion queues, generating flow metadata for different flows of packets representing usage of at least one resource of the node by each of the different flows of packets initiated by at least one user space application,

wherein at least one of one or more respective memory request rates and one or more respective resource allocations for the at least one user space application is based at least in part on the generated flow metadata.

20. The node of claim 18 , wherein the kernel space means is further for executing a VS kernel module in the kernel space, the VS kernel module configured to, based on the memory request rate, add at least one of a congestion indicator and the memory request rate to one or more packets of at least one flow of packets.

Assignments (3)
PATENT COLLATERAL AGREEMENT - A&R LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064715/0001 →
PATENT COLLATERAL AGREEMENT - DDTL LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 067045/0156 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: RADI, MARJAN; VUCINIC, DEJAN
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 060155/0445 →
Continuity (1)
Related Publication 20230401079A1 · Dec 14, 2023
References Cited (109)
US 6381686B1 · Imamura · 2002 [cited by applicant]
US 8412907B1 · Dunshea et al. · 2013 [cited by applicant]
US 8700727B1 · Gole et al. · 2014 [cited by applicant]
US 10027697B1 · Babun et al. · 2018 [cited by applicant]
US 10362149B2 · Biederman et al. · 2019 [cited by applicant]
US 10530711B2 · Yu et al. · 2020 [cited by applicant]
US 10628560B1 · Siranni et al. · 2020 [cited by applicant]
US 10706147B1 · Pohlack · 2020 [cited by applicant]
US 10754707B2 · Tamir et al. · 2020 [cited by applicant]
US 10757021B2 · Man et al. · 2020 [cited by applicant]
US 11134025B2 · Billore et al. · 2021 [cited by applicant]
US 11223579B2 · Lu · 2022 [cited by applicant]
US 20020143843A1 · Mehta · 2002 [cited by applicant]
US 20050257263A1 · Keohane et al. · 2005 [cited by applicant]
US 20060101466A1 · Kawachiya et al. · 2006 [cited by applicant]
US 20070067840A1 · Young · 2007 [cited by applicant]
US 20090249357A1 · Chanda et al. · 2009 [cited by applicant]
US 20100161976A1 · Bacher · 2010 [cited by applicant]
US 20110126269A1 · Youngworth · 2011 [cited by examiner]
US 20120198192A1 · Balasubramanian et al. · 2012 [cited by applicant]
US 20120207026A1 · Sato · 2012 [cited by applicant]
US 20140143365A1 · Guerin et al. · 2014 [cited by applicant]
US 20140283058A1 · Gupta · 2014 [cited by applicant]
US 20150006663A1 · Huang · 2015 [cited by applicant]
US 20150319237A1 · Hussain et al. · 2015 [cited by applicant]
US 20150331622A1 · Chiu et al. · 2015 [cited by applicant]
US 20170163479A1 · Wang et al. · 2017 [cited by applicant]
US 20170269991A1 · Bazarsky et al. · 2017 [cited by applicant]
US 20180004456A1 · Talwar et al. · 2018 [cited by applicant]
US 20180012020A1 · Prvulovic et al. · 2018 [cited by applicant]
US 20180032423A1 · Bull et al. · 2018 [cited by applicant]
US 20180060136A1 · Herdrich et al. · 2018 [cited by applicant]
US 20180173555A1 · Lutas · 2018 [cited by applicant]
US 20180191632A1 · Biederman et al. · 2018 [cited by applicant]
US 20180341419A1 · Wang et al. · 2018 [cited by applicant]
US 20180357176A1 · Wang · 2018 [cited by applicant]
US 20190227936A1 · Jang · 2019 [cited by applicant]
US 20190280964A1 · Michael et al. · 2019 [cited by applicant]
US 20200034538A1 · Woodward et al. · 2020 [cited by applicant]
US 20200201775A1 · Zhang et al. · 2020 [cited by applicant]
US 20200274952A1 · Waskiewicz et al. · 2020 [cited by applicant]
US 20200285591A1 · Luo et al. · 2020 [cited by applicant]
US 20200322287A1 · Connor et al. · 2020 [cited by applicant]
US 20200403905A1 · Allen et al. · 2020 [cited by applicant]
US 20200409821A1 · Terada et al. · 2020 [cited by applicant]
US 20210019197A1 · Tamir et al. · 2021 [cited by applicant]
US 20210058424A1 · Chang et al. · 2021 [cited by applicant]
US 20210103505A1 · Tsuchiya · 2021 [cited by examiner]
US 20210149763A1 · Ranganathan et al. · 2021 [cited by applicant]
US 20210157740A1 · Benhanokh et al. · 2021 [cited by applicant]
US 20210240621A1 · Fu et al. · 2021 [cited by applicant]
US 20210266253A1 · He et al. · 2021 [cited by applicant]
US 20210320881A1 · Coyle et al. · 2021 [cited by applicant]
US 20210377150A1 · Dugast et al. · 2021 [cited by applicant]
US 20220035698A1 · Vankamamidi et al. · 2022 [cited by applicant]
US 20220121362A1 · Liu et al. · 2022 [cited by applicant]
US 20220294883A1 · Pope et al. · 2022 [cited by applicant]
US 20220350516A1 · Bono et al. · 2022 [cited by applicant]
US 20220357886A1 · Pitchumani et al. · 2022 [cited by applicant]
US 20220414968A1 · Wiegert et al. · 2022 [cited by applicant]
CN 106603409A · 2017 [cited by applicant]
CN 112351250A · 2021 [cited by applicant]
EP 3358456A1 · 2018 [cited by applicant]
EP 3598309B1 · 2022 [cited by applicant]
KR 1020190090331A · 2019 [cited by applicant]
WO 2018086569A1 · 2018 [cited by applicant]
WO 2018145725A1 · 2018 [cited by applicant]
WO 2021226948A1 · 2021 [cited by applicant]
Maefeichen.com; “Setup the extended Berkeley Packet Filter (eBPF) Environment”; Maofei's Blog; Dec. 9, 2021; available at: https://maofeichen.com/setup-the-extended-berkeley-packet-filter-ebpf-environment/. [cited by applicant]
International Search Report and Written Opinion dated Oct. 25, 2022 from International Application No. PCT/ US2022/030414, 11 pages. [cited by applicant]
Sabella et al.; “Using eBPF for network traffic analysis”; available at: Year: 2018; https://www.ntop.org/wp-content/uploads/2018/10/Sabella.pdf. [cited by applicant]
Bachl et al.; “A flow-based IDS using Machine Learning in EBPF”; Cornell University; Feb. 19, 2021; available at https://arxiv.org/abs/2102.09980. [cited by applicant]
Caviglione et al.; “Kernel-level tracing for detecting stegomalware and covert channels in Linux environments”; Computer Networks 191; Mar. 2021; available at: https://www.researchgate.net/publication/350182568_Kernel-l… [cited by applicant]
Dimolianis et al.; “Signature-Based Traffic Classification and Mitigation for DDOS Attacks Using Programmable Network Data Planes”; IEEE Access; Jul. 7, 2021; available at: https://ieeexplore.ieee.org/stamp/stamp.jsp?am… [cited by applicant]
Jun Li; “Efficient Erasure Coding In Distributed Storage Systems”; A thesis submitted in conformity with the requirements for the degree of Doctor of Philosophy Graduate Department of Electrical and Computer Engineering… [cited by applicant]
Lakshmi J. Mohan; “Erasure codes for optimal performance in geographically distributed storage systems”; Apr. 2018; School of Computing and Information Systems, University of Melbourne; available at: https://minerva-acc… [cited by applicant]
Navarre et al.; “SRv6-FEC: Bringing Forward Erasure Correction to IPv6 Segment Routing”; SIGCOMM '21: Proceedings of the SIGCOMM '21 Poster and Demo Sessions; Aug. 2021; pp. 45-47; available at: https://dl.acm.org/doi/1… [cited by applicant]
Van Schaik et al.; “RIDL: Rogue In-Flight Data Load”; Proceedings—IEEE Symposium on Security and Privacy; May 2019; available at: https://mdsattacks.com/files/ridl.pdf. [cited by applicant]
Xhonneux et al.; “Flexible failure detection and fast reroute using eBPF and SRv6”; 2018 14th International Conference on Network and Service Management (CNSM); Nov. 2018; available at: https://dl.ifip.org/db/conf/cnsm/… [cited by applicant]
Zhong et al.; “Revisiting Swapping in User-space with Lightweight Threading”; arXiv:2107.13848v1; Jul. 29, 2021; available at: https://deepai.org/publication/revisiting-swapping-in-user-space-with-lightweight-threading. [cited by applicant]
Baidya et al.; “eBPF-based Content and Computation-aware Communication for Real-time Edge Computing”; IEEE International Conference on Computer Communications (INFOCOM Workshops); May 8, 2018; available at https://arxiv… [cited by applicant]
Barbalace et al.; “blockNDP: Block-storage Near Data Processing”; University of Edinburgh, Huawei Dresden Research Center, Huawei Munich Research Center, TUM; Dec. 2020; 8 pages; available at https://dl.acm.org/doi/10.1… [cited by applicant]
Blin et al.; “Toward an in-kernel high performance key-value store implementation”; Oct. 2019; 38th Symposium on Reliable Distributed Systems (SRDS); available at: https://ieeexplore.ieee.org/document/9049596. [cited by applicant]
Enberg et al.; “Partition-Aware Packet Steering Using XDP and eBPF for Improving Application-Level Parallelism”; ENCP; Dec. 9, 2019; 7 pages; available at: https://penberg.org/papers/xdp-steering-encp19.pdf. [cited by applicant]
Kicinski et al.; “eBPF Hardware Offload to SmartNICs: cls_bpf and XDP”; Netronome Systems Cambridge, United Kingdom; 2016; 6 pages; available at https://www.netronome.com/media/documents/eBPF_HW_OFFLOAD_HNiMne8_2_.pdf. [cited by applicant]
Kourtis et al.; “Safe and Efficient Remote Application Code Execution on Disaggregated NVM Storage with eBPF”; Feb. 25, 2020; 8 pages; available at https://arxiv.org/abs/2002.11528. [cited by applicant]
Wu et al.; “BPF for storage: an exokernel-inspired approach”; Columbia University, University of Utah, VMware Research; Feb. 25, 2021; 8 pages; available at: https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s0… [cited by applicant]
Pending U.S. Appl. No. 17/561,898, filed Dec. 24, 2021, entitled “In-Kernel Caching for Distributed Cache”, Marjan Radi. [cited by applicant]
Pending U.S. Appl. No. 17/571,922, filed Jan. 10, 2022, entitled “Computational Acceleration for Distributed Cache”, Marjan Radi. [cited by applicant]
Pending U.S. Appl. No. 17/665,330, filed Feb. 4, 2022, entitled “Error Detection and Data Recovery for Distributed Cache”, Marjan Radi. [cited by applicant]
Pending U.S. Appl. No. 17/683,737, filed Mar. 1, 2022, entitled “Detection of Malicious Operations for Distributed Cache”, Marjan Radi. [cited by applicant]
Pending U.S. Appl. No. 17/741,244, filed May 10, 2022, entitled “In-Kernel Cache Request Queuing for Distributed Cache”, Marjan Radi. [cited by applicant]
Tu et al.; “Bringing the Power of eBPF to Open vSwitch”; Linux Plumber 2018; available at: http://vger.kernel.org/lpc_net2018_talks/ovs-ebpf-afxdp.pdf. [cited by applicant]
Pending U.S. Appl. No. 17/829,712, filed Jun. 1, 2022, entitled “Context-Aware NVMe Processing in Virtualized Environments”, Marjan Radi. [cited by applicant]
International Search Report and Written Opinion dated Nov. 18, 2022 from International Application No. PCT/US2022/030437, 10 pages. [cited by applicant]
Kang et al.; “Enabling Cost-effective Data Processing with Smart SSD”; 2013 IEEE 29th Symposium on Mass Storage Systems and Technologies (MSST); available at: https://pages.cs.wisc.edu/˜yxy/cs839-s20/papers/SmartSSD2.pd… [cited by applicant]
Bijlani et al.; “Extension Framework for File Systems in User space”; Jul. 2019; Usenix; available at: https://www.usenix.org/conference/atc19/presentation/bijlani. [cited by applicant]
Brad Fitzpatrick; “Distributed Caching with Memcached”; Aug. 1, 2004; Linux Journal; available at: https://www.linuxjournal.com/article/7451. [cited by applicant]
Roderick W. Smith; “The Definitive Guide to Samba 3”; 2004; APress Media; pp. 332-336; available at: https://link.springer.com/book/10.1007/978-1-4302-0683-5. [cited by applicant]
Wu et al.; “NCA: Accelerating Network Caching with express Data Path”; Nov. 2021; IEEE; available at https://ieeexplore.ieee.org/abstract/document/9680837. [cited by applicant]
Patterson et al.; “Computer Architecture: A Quantitative Approach”; 1996; Morgan Kaufmann; 2nd ed.; pp. 378-380. [cited by applicant]
Pending U.S. Appl. No. 17/850,767, filed Jun. 27, 2022, entitled “Memory Coherence in Virtualized Environments”, Marjan Radi. [cited by applicant]
Gao et al.; “OVS-CAB: Efficient rule-caching for Open vSwitch hardware offloading”; Computer Networks; Apr. 2021; available at:https://www.sciencedirect.com/science/article/abs/pii/S1389128621000244. [cited by applicant]
Pfaff et al.; “The Design and Implementation of Open vSwitch”; Usenix; May 4, 2015; available at: https://www.usenix.org/conference/nsdi15/technical-sessions/presentation/pfaff. [cited by applicant]
Ghigoff et al., “BMC: Accelerating Memcached using Safe In-kernel Caching and Pre-stack Processing”; In: 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 2021); p. 487-501; Apr. 14, 2021. [cited by applicant]
International Search Report and Written Opinion dated Sep. 30, 2022 from International Application No. PCT/US2022/029527, 9 pages. [cited by applicant]
Anderson et al.; “Assise: Performance and Availability via Client-local NVM in a Distributed File System”; the 14th USENIX Symposium on Operating Systems Design and Implementation; Nov. 6, 2020; available at: https://ww… [cited by applicant]
Pinto et al.; “Hoard: A Distributed Data Caching System to Accelerate Deep Learning Training on the Cloud”; arXiv; Dec. 3, 2018; available at: https://arxiv.org/pdf/1812.00669.pdf. [cited by applicant]
International Search Report and Written Opinion dated Oct. 7, 2022 from International Application No. PCT/US2022/030044, 10 pages. [cited by applicant]