IP Library › Granted Patent US 12,314,178
Granted Patent B2
US 12,314,178 · App. 17/344,763 · Granted May 27, 2025

Management of distributed shared memory

Inventors: Kshitij A. Doshi (Tempe, AZ); Francesc Guim Bernat (Barcelona, ES); Suraj Prabhakaran (Aachen, DE); Tushar Sudhakar Gohad (Phoenix, AZ)
Assignee: Intel Corporation
G06F12/0813G06F12/1009H04L41/0893
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,178
App. No.
17/344,763
Filed
Jun 10, 2021
Granted
May 27, 2025
Kind
B2
Art Unit
2136
USPC
711/206
Abstract

Examples described herein relate to a network interface device. In some examples, the network interface device includes a device interface; input/output circuitry to receive Ethernet compliant packets and output Ethernet compliant packets; circuitry to monitor a particular page for a rate of data copying among nodes within a group of two or more nodes; and circuitry to perform one or more actions based, at least in part, on the rate of data copying among the nodes within the group of two or more nodes to attempt to reduce a number of copy operations of the data among the nodes within the group of two or more nodes, wherein the group of two or more nodes are part of a distributed shared memory (DSM).

Claims (29)

1. An apparatus comprising:

a network interface device comprising:

a device interface;

input/output circuitry to receive Ethernet compliant packets and transmit Ethernet compliant packets;

circuitry to monitor copying of a particular page of data among nodes within a group of two or more nodes based on heatmap data indicative of copying activity of the particular page of data among the group of two or more nodes; and

circuitry to perform one or more actions based, at least in part, on the data copying activity among the nodes within the group of two or more nodes based on the heatmap data to attempt to reduce a number of copy operations of the particular page of data among the nodes within the group of two or more nodes, wherein the group of two or more nodes are part of a distributed shared memory (DSM) and wherein the one or more actions comprise migration of an accessor process that is to access the particular page of data to execute on a target node within the group of two or more nodes and wherein the target node is to store the particular page of data accessed by the accessor process of the data.

2. The apparatus of claim 1 , wherein the one or more actions comprise create multiple replicas of the particular page of data and defer merge of the particular page of data.

3. The apparatus of claim 1 , wherein the one or more actions comprise split the particular page of data into smaller ranges to reduce a size of data copied.

4. The apparatus of claim 1 , wherein the one or more actions comprise selection of a coordinator node to manage one or more updates to the particular page of data and to provide an up-to-date copy of the data in the particular page.

5. The apparatus of claim 1 , wherein the network interface device is part of a cluster of network devices wherein applications execute on multiple processors of the network devices and access data from logically shared memory and wherein the accessed data is physically distributed across memory devices over a scale-out network.

6. The apparatus of claim 1 , wherein the network interface device comprises one or more of: a network interface controller (NIC), a SmartNIC, infrastructure processing unit (IPU), switch, and/or switch with programmable packet processing pipeline.

7. The apparatus of claim 1 , comprising a host node coupled to the device interface, wherein the host node comprises:

at least one memory device and

at least one processor to execute an application that is to access the particular page of data stored within the group of two or more nodes, wherein the group of two or more nodes are consistent with a DSM model.

8. The apparatus of claim 1 , wherein the network interface device is to receive the heat map data in one or more packets.

9. A method comprising:

a network interface device monitoring a particular address range for a rate of data copying among nodes within a group of two or more nodes based on heatmap data indicative of copying activity of the particular address range of data among the group of two or more nodes, wherein the network interface device comprises: a network interface to receive and transmit packets and a host interface and

the network interface device performing one or more actions based, at least in part, on the rate of data copying among the nodes within the group of two or more nodes, wherein the one or more actions comprise migration of an accessor to execute on a target node within the group of two or more nodes and wherein, after migration of the accessor to execute on the target node, the target node stores data associated with the particular address range accessed by the accessor and the target node executes the accessor.

10. The method of claim 9 , wherein the one or more actions comprise copy the data to a node that is a fewer number of hops away from a node that accesses the data.

11. The method of claim 9 , wherein the one or more actions comprise split the particular address range of data into smaller ranges to reduce a transmitted size of data associated with the particular address range.

12. The method of claim 9 , wherein the one or more actions comprise selection of a coordinator node to manage one or more updates to data stored in the particular address range and to provide an up-to-date copy of the data in the particular address range.

13. The method of claim 9 , wherein the network interface device is part of a cluster of network devices wherein applications execute on multiple processors of the network devices and access data from logically shared memory and wherein the accessed data is physically distributed across memory devices over a scale-out network.

14. At least one non-transitory computer-readable medium comprising instructions stored thereon, that if executed by one or more processors, cause the one or more processors to:

configure a network interface device to monitor a particular address range for a rate of data copying among nodes within a group of two or more nodes based on heatmap data indicative of the rate of data copying activity of the particular address range of data among the group of two or more nodes, wherein the network interface device comprises: a network interface to receive and transmit packets and a host interface and

configure the network interface device to perform one or more actions based, at least in part, on the rate of data copying among the nodes within the group of two or more nodes, wherein the one or more actions comprise migration of an application to execute on a target node within the group of two or more nodes and wherein, after migration of the application to execute on the target node, the target node stores data associated with the particular address range accessed by the application and the target node executes the application and wherein the application comprises one or more of: a virtual machine (VM), container, service, microservice, or executable binary.

15. The computer-readable medium of claim 14 , wherein the one or more actions comprise copy the data to a node that is a fewer number of hops away from a node that accesses the data.

16. The computer-readable medium of claim 14 , wherein the one or more actions comprise split the particular address range of data into smaller ranges to reduce a transmitted size of data associated with the particular address range.

17. The computer-readable medium of claim 14 , wherein the one or more actions comprise selection of a coordinator node to manage one or more updates to the data stored in the particular address range and to provide a true copy of the data in the particular address range.

18. The computer-readable medium of claim 14 , wherein an orchestrator, driver, and/or operating system (OS) is to configure the network interface device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2025
From: DOSHI, KSHITIJ A.; GUIM BERNAT, FRANCESC; GOHAD, TUSHAR SUDHAKAR
To: INTEL CORPORATION
Reel/Frame 070816/0419 →
Continuity (2)
Provisional Application 63130664 · Dec 26, 2020
Related Publication 20210303477A1 · Sep 30, 2021
References Cited (32)
US 5822773A · Pritchard · 1998 [cited by examiner]
US 7949737B2 · Tan · 2011 [cited by examiner]
US 20010047460A1 · Kobayashi · 2001 [cited by examiner]
US 20020144058A1 · Burger · 2002 [cited by examiner]
US 20090257386A1 · Achir · 2009 [cited by examiner]
US 20120272029A1 · Zhang et al. · 2012 [cited by applicant]
US 20140082308A1 · Naruse · 2014 [cited by examiner]
US 20140136773A1 · Michalak · 2014 [cited by applicant]
US 20140156777A1 · Subbiah et al. · 2014 [cited by applicant]
US 20150370702A1 · Voigt · 2015 [cited by examiner]
US 20180006923A1 · Gao et al. · 2018 [cited by applicant]
US 20180365167A1 · Eckert · 2018 [cited by examiner]
US 20200349074A1 · Kucherov · 2020 [cited by examiner]
US 20220091738A1 · Patil et al. · 2022 [cited by applicant]
US 20220164118A1 · Hsu · 2022 [cited by examiner]
EP 0308047A2 · 1989 [cited by examiner]
KR 1020130048594A · 2013 [cited by applicant]
WO WO2016070302A1 · 2016 [cited by examiner]
WO WO2020231467A1 · 2020 [cited by examiner]
H. Jonathan Chao; Bin Liu, “HighSpeed Router Chip Set,” in High Performance Switches and Routers , IEEE, 2007, pp. 538-605, ch16. [cited by examiner]
C. Scheurich and M. Dubois, “Dynamic page migration in multiprocessors with distributed global memory,” in IEEE Transactions on Computers, vol. 38, No. 8, pp. 1154-1163, Aug. 1989. [cited by examiner]
D. S. Nikolopoulos, T. S. Papatheodorou, C. D. Polychronopoulos, J. Labarta and E. Ayguade, “Is Data Distribution Necessary in OpenMP?,” SC '00: Proceedings of the 2000 ACM/IEEE Conference on Supercomputing, Dallas, TX,… [cited by examiner]
Amza, Cristiana, et al., “TreadMarks: Shared Memory Computing on Networks of Work-stations”, IEEE Computer, Feb. 1995, 26 pages. [cited by applicant]
Ibel, Maximilian et al., “High-Performance Cluster Computing Using SCI. Hot Interconnects”, Aug. 1997 13 pages. [cited by applicant]
Itzkovitz, Ayal et al., “Millipede: a User-Level NT-Based Distributed Shared Memory System with Thread Migration and Dynamic Run-Time Optimization of Memory References”, USENIX Windows NT Workshop, Aug. 1997, 2 pages. [cited by applicant]
Jiang, Dave, “Introducing the Intel® Data Streaming Accelerator (Intel® DSA)”, https://01.org/blogs/2019/introducing-intel-data-streaming-accelerator, Nov. 2019, 5 pages. [cited by applicant]
Keleher, Pete et al., “TreadMarks: Distributed Shared Memory on Standard Workstations and Operating Systems”, USENIX Winter 1994, 17 pages. [cited by applicant]
Koch, Povl T. et al., “Global Management of Coherent Shared Memory on an SCI Cluster”, https://www.researchgate.net/publication/2578306_Global_Management_of_Coherent_Shared_Memory_on_an_SCI_Cluster, Dec. 1999, 7 pages. [cited by applicant]
Kumar, Akhilesh, “New Intel® Mesh Architecture: The ‘Superhighway’ of the Data Center”, Intel, White Paper, Apr. 2020, 3 pages. [cited by applicant]
Paas, Sven M. et al., “Computing on a Cluster of PCs: Project Overview and Early Experiences”, CSR-97-05 Chemnitzer Informatik-Berichte, Winter 1997. 13 pages. [cited by applicant]
Speight, Evan and Bennett, John K., “Brazos: A Third Generation DSM System”, First USENIX Windows NT Workshop, Aug. 1997, 13 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US21/51794, Mailed Jan. 3, 2022, 8 pages. [cited by applicant]