IP Library › Granted Patent US 8,805,952
Granted Patent B2
US 8,805,952 · App. 13/343,239 · Granted Aug 12, 2014

Administering globally accessible memory space in a distributed computing system

Inventors: Tsai-Yang Jea (Poughkeepsie, NY); Yuan Yuan Nie (Beijing, CN)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,805,952
App. No.
13/343,239
Granted
Aug 12, 2014
Kind
B2
Abstract

In a distributed computing system that includes compute nodes that include computer memory, globally accessible memory space is administered by: for each compute node: mapping a memory region of a predefined size beginning at a predefined address; executing one or more memory management operations within the memory region, including, for each memory management operation executed within the memory region: executing the operation collectively by all compute nodes, where the operation includes a specification of one or more parameters and the parameters are the same across all compute nodes; receiving, by each compute node from a deterministic memory management module in response to the memory management operation, a return value, where the return value is the same across all compute nodes; entering, by each compute node after local completion of the memory management operation, a barrier; and when all compute nodes have entered the barrier, resuming execution.

Claims (38)

1. An apparatus for administering globally accessible memory space in a distributed computing system, the distributed computing system comprising a plurality of compute nodes coupled for data communications by one or more data communications networks, each compute node comprising computer memory, the apparatus comprising a computer processor, a computer memory operatively coupled to the computer processor, the computer memory having stored thereon computer program instructions that, when executed, cause the computer processor to carry out a method steps of:

for each compute node:

mapping a memory region of a predefined size beginning at a predefined address;

executing one or more memory management operations within the memory region, including, for each memory management operation executed within the memory region:

executing the operation collectively by all compute nodes, where the operation includes a specification of one or more parameters and the parameters are the same across all compute nodes;

reallocating, by each compute node in a collective operation, a buffer;

receiving, by each compute node from a deterministic memory management module in response to the memory management operation, a return value, where the return value is the same across all compute nodes, wherein receiving the return value from the deterministic memory management module further comprises receiving a local memory address of the reallocated buffer within the memory region, where the local memory address is the same address on all compute nodes;

entering, by each compute node after local completion of the memory management operation, a barrier; and

when all compute nodes have entered the barrier, resuming execution.

2. The apparatus of claim 1 wherein:

executing one or more memory management operations further comprises allocating, by each compute node in a collective operation, a buffer; and

receiving the return value from the deterministic memory management module further comprises receiving a local memory address of the allocated buffer within the memory region, where the local memory address is the same address on all compute nodes.

3. The apparatus of claim 2 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of:

storing, by a first compute node, data in the allocated buffer of a second compute node through use of the first compute node's local memory address.

4. The apparatus of claim 2 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of:

retrieving, by the second compute node, data from the first compute node's allocated buffer through use of the second compute node's local memory address.

5. The apparatus of claim 1 wherein:

executing one or more memory management operations further comprises deallocating, by each compute node in a collective operation, a buffer; and

receiving the return value from the deterministic memory management module further comprises receiving a result code, where the result code is the same address on all compute nodes.

6. A computer program product for administering globally accessible memory space in a distributed computing system, the distributed computing system comprising a plurality of compute nodes coupled for data communications by one or more data communications networks, each compute node comprising computer memory, the computer program product stored on a computer readable storage medium, wherein the computer readable storage medium is not a signal, the computer program product comprising computer program instructions that, when executed, cause a computer to carry out a method steps of:

for each compute node:

mapping a memory region of a predefined size beginning at a predefined address;

executing one or more memory management operations within the memory region, including, for each memory management operation executed within the memory region:

executing the operation collectively by all compute nodes, where the operation includes a specification of one or more parameters and the parameters are the same across all compute nodes;

reallocating, by each compute node in a collective operation, a buffer;

receiving, by each compute node from a deterministic memory management module in response to the memory management operation, a return value, where the return value is the same across all compute nodes, wherein receiving the return value from the deterministic memory management module further comprises receiving a local memory address of the reallocated buffer within the memory region, where the local memory address is the same address on all compute nodes;

entering, by each compute node after local completion of the memory management operation, a barrier; and

when all compute nodes have entered the barrier, resuming execution.

7. The computer program product of claim 6 wherein:

executing one or more memory management operations further comprises allocating, by each compute node in a collective operation, a buffer; and

receiving the return value from the deterministic memory management module further comprises receiving a local memory address of the allocated buffer within the memory region, where the local memory address is the same address on all compute nodes.

8. The computer program product of claim 7 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of:

storing, by a first compute node, data in the allocated buffer of a second compute node through use of the first compute node's local memory address.

9. The computer program product of claim 7 further comprising computer program instructions that, when executed by the computer processor, cause the apparatus to carry out the step of:

retrieving, by the second compute node, data from the first compute node's allocated buffer through use of the second compute node's local memory address.

10. The computer program product of claim 6 wherein:

executing one or more memory management operations further comprises deallocating, by each compute node in a collective operation, a buffer; and

receiving the return value from the deterministic memory management module further comprises receiving a result code, where the result code is the same address on all compute nodes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2012
From: JEA, TSAI-YANG; NIE, YUAN YUAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 027476/0544 →
Continuity (1)
Related Publication 20130173738A1 · Jul 4, 2013