IP Library › Granted Patent US 9,170,864
Granted Patent B2
US 9,170,864 · App. 12/362,137 · Granted Oct 27, 2015

Data processing in a hybrid computing environment

Inventors: Charles J. Archer (Rochester, MN); James E. Carey (Rochester, MN); Matthew W. Markland (Rochester, MN); Philip J. Sanders (Rochester, MN); Timothy J. Schimke (Stewartville, MN)
Assignee: International Business Machines Corporation
G06F9/544G06F9/5055G06F9/5066G06F2209/509G06F2209/5017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,170,864
App. No.
12/362,137
Granted
Oct 27, 2015
Kind
B2
Abstract

Data processing in a hybrid computing environment that includes a host computer, a plurality of accelerators, the host computer and the accelerators adapted to one another for data communications by a system level message passing module, the host computer having local memory shared remotely with the accelerators, the accelerators having local memory for the plurality of accelerators shared remotely with the host computer, where data processing according to embodiments of the present invention includes performing, by the plurality of accelerators, a local reduction operation with the local shared memory for the accelerators; writing remotely, by one of the plurality of accelerators to the shared memory local to the host computer, a result of the local reduction operation; and reading, by the host computer from shared memory local to the host computer, the result of the local reduction operation.

Claims (36)

1. A method of data processing in a hybrid computing environment, the hybrid computing environment comprising a host computer having a host computer architecture, a plurality of accelerators having an accelerator architecture, wherein each of the plurality of accelerators has a dedicated portion of the local shared memory for the accelerators, the accelerator architecture optimized, with respect to the host computer architecture, for speed of execution of a particular class of computing functions, the host computer and the accelerators adapted to one another for data communications by a system level message passing module, the host computer having local memory shared remotely with the accelerators, the accelerators having local memory for the plurality of accelerators shared remotely with the host computer, the method comprising:

assigning to each accelerator a rank in a logical tree and a specific function to perform, the results of the specific function to be stored in the accelerator's dedicated portion;

performing, by the plurality of accelerators, a local reduction operation with the local shared memory for the accelerators, wherein performing the local reduction operation with the location shared memory for the accelerators comprises:

performing, by each accelerator, the accelerator's assigned specific function including:

determining, by the accelerator in dependence upon the accelerator's assigned rank, whether the accelerator is authorized to perform the accelerator's assigned specific function; and

if the accelerator is authorized to perform the accelerator's assigned specific function, performing, by the accelerator, the specific function and incrementing, by the accelerator, a counter; and

storing locally by each accelerator the results of the accelerator's assigned specific function in the accelerator's dedicated portion of the local shared memory for the accelerators;

writing remotely, by one of the plurality of accelerators to the shared memory local to the host computer, a result of the local reduction operation; and

reading, by the host computer from shared memory local to the host computer, the result of the local reduction operation.

2. The method of claim 1 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having no children in the logical tree to store the accelerator's contribution data in the accelerator's dedicated portion of the local shared memory for the accelerators.

3. The method of claim 1 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having one or more children in the logical tree to read the contents of the children's dedicated portions of the local shared memory for the accelerators and to perform the specific function with the read contents and contribution data of the accelerator having one or more children.

4. The method of claim 1 wherein determining whether the accelerator is authorized to perform the accelerator's assigned specific function further comprises determining whether the counter exceeds a predetermined threshold for a level of depth of the logic tree to which the accelerator is assigned.

5. A hybrid computing environment for data processing, the hybrid computing environment comprising a host computer having a host computer architecture, a plurality of accelerators having an accelerator architecture, wherein each of the plurality of accelerators has a dedicated portion of the local shared memory for the accelerators, the accelerator architecture optimized, with respect to the host computer architecture, for speed of execution of a particular class of computing functions, the host computer and the accelerators adapted to one another for data communications by a system level message passing module, the host computer having local memory shared remotely with the accelerators, the accelerators having local memory for the plurality of accelerators shared remotely with the host computer, the plurality of accelerators comprising computer program instructions capable of:

assigning to each accelerator a rank in a logical tree and a specific function to perform, the results of the specific function to be stored in the accelerator's dedicated portion;

performing, by the plurality of accelerators, a local reduction operation with the local shared memory for the accelerators, wherein performing the local reduction operation with the location shared memory for the accelerators comprises:

performing, by each accelerator, the accelerator's assigned specific function including:

determining, by the accelerator in dependence upon the accelerator's assigned rank, whether the accelerator is authorized to perform the accelerator's assigned specific function; and

if the accelerator is authorized to perform the accelerator's assigned specific function, performing, by the accelerator, the specific function and incrementing, by the accelerator, a counter; and

storing locally by each accelerator the results of the accelerator's assigned specific function in the accelerator's dedicated portion of the local shared memory for the accelerators;

writing remotely, by one of the plurality of accelerators to the shared memory local to the host computer, a result of the local reduction operation; and

the host computer comprising computer program instructions capable of reading, by the host computer from shared memory local to the host computer, the result of the local reduction operation.

6. The hybrid computing environment of claim 5 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having no children in the logical tree to store the accelerator's contribution data in the accelerator's dedicated portion of the local shared memory for the accelerators.

7. The hybrid computing environment of claim 5 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having one or more children in the logical tree to read the contents of the children's dedicated portions of the local shared memory for the accelerators and to perform the specific function with the read contents and contribution data of the accelerator having one or more children.

8. The hybrid computing environment of claim 5 wherein determining whether the accelerator is authorized to perform the accelerator's assigned specific function further comprises determining whether the counter exceeds a predetermined threshold for a level of depth of the logic tree to which the accelerator is assigned.

9. A computer program product for data processing in a hybrid computing environment, the hybrid computing environment comprising a host computer having a host computer architecture, a plurality of accelerators having an accelerator architecture, wherein each of the plurality of accelerators has a dedicated portion of the local shared memory for the accelerators, the accelerator architecture optimized, with respect to the host computer architecture, for speed of execution of a particular class of computing functions, the host computer and the accelerators adapted to one another for data communications by a system level message passing module, the host computer having local memory shared remotely with the accelerators, the accelerators having local memory for the plurality of accelerators shared remotely with the host computer, the computer program product disposed in a computer readable, recordable storage medium, the computer program product comprising computer program instructions capable of:

assigning to each accelerator a rank in a logical tree and a specific function to perform, the results of the specific function to be stored in the accelerator's dedicated portion;

performing, by the plurality of accelerators, a local reduction operation with the local shared memory for the accelerators, wherein performing the local reduction operation with the location shared memory for the accelerators comprises:

performing, by each accelerator, the accelerator's assigned specific function including:

determining, by the accelerator in dependence upon the accelerator's assigned rank, whether the accelerator is authorized to perform the accelerator's assigned specific function; and

if the accelerator is authorized to perform the accelerator's assigned specific function, performing, by the accelerator, the specific function and incrementing, by the accelerator, a counter; and

storing locally by each accelerator the results of the accelerator's assigned specific function in the accelerator's dedicated portion of the local shared memory for the accelerators;

writing remotely, by one of the plurality of accelerators to the shared memory local to the host computer, a result of the local reduction operation; and

reading, by the host computer from shared memory local to the host computer, the result of the local reduction operation.

10. The computer program product of claim 9 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having no children in the logical tree to store the accelerator's contribution data in the accelerator's dedicated portion of the local shared memory for the accelerators.

11. The computer program product of claim 9 wherein assigning each accelerator a specific function to perform further comprises instructing an accelerator having one or more children in the logical tree to read the contents of the children's dedicated portions of the local shared memory for the accelerators and to perform the specific function with the read contents and contribution data of the accelerator having one or more children.

12. The computer program product of claim 9 wherein determining whether the accelerator is authorized to perform the accelerator's assigned specific function further comprises determining whether the counter exceeds a predetermined threshold for a level of depth of the logic tree to which the accelerator is assigned.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2009
From: ARCHER, CHARLES J.; CAREY, JANES E.; MARKLAND, MATTHEW W.; SANDERS, PHILIP J.; SCHIMKE, TIMOTHY J.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 022175/0628 →
Continuity (1)
Related Publication 20100191823A1 · Jul 29, 2010