IP Library Granted Patent US 7,647,484
Granted Patent B2
US 7,647,484 · App. 11/678,208 · Granted Jan 12, 2010

Low-impact performance sampling within a massively parallel computer

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,647,484
App. No.
11/678,208
Granted
Jan 12, 2010
Kind
B2
Abstract

An apparatus, program product and method sample at different times nodes that are performing similar work. Performance data associated with first and second node subsets performing the similar work are sampled at different times, e.g., in a round-robin fashion, and in accordance with a given sampling rate. The performance data is analyzed. Nodes whose performance suffers as a result of a sampling operation may be identified and removed from a subsequent operation.

Claims (25)

1. A method of sampling performance data of a plurality of interconnected nodes comprising a massively parallel computer system, the method comprising:

identifying a plurality of nodes doing similar work;

distributing an overhead for performance sampling across the plurality of nodes by sampling at different times performance data associated with first and second subsets of the plurality of nodes, wherein distributing the overhead for performance sampling across the plurality of nodes includes sequentially sampling the plurality of nodes in a round-robin fashion; and

analyzing the performance data.

2. The method of claim 1 , wherein sampling the performance data of the first subset further comprises sampling the performance data of a first node of the plurality of nodes.

3. The method of claim 1 , wherein identifying the plurality of nodes doing the similar work further comprises identifying a plurality of nodes executing a common application.

4. The method of claim 1 further comprising adjusting a sampling rate at which the performance data is sampled.

5. The method of claim 1 , wherein identifying the plurality of nodes doing the similar work further comprises automatically identifying the plurality of nodes doing the similar work.

6. The method of claim 1 , wherein sampling the performance data further comprises sampling the first and second subsets at a common sampling rate.

7. The method of claim 1 further comprising determining that the first subset is performing unsatisfactorily.

8. The method of claim 7 further comprising removing the first subset from a subsequent sampling operation in response to determining that the first subset is performing unsatisfactorily.

9. An apparatus, comprising:

a hardware processor; and

program code configured to be executed by the processor to sample performance data of a plurality of interconnected nodes comprising a massively parallel computer system, the program code configured to identify a plurality of nodes doing similar work, to distribute an overhead for performance sampling across the plurality of nodes by sampling at different times performance data associated with first and second subsets of the plurality of nodes, and to analyze the performance data, wherein the program code distributes the overhead for performance sampling across the plurality of nodes by sequentially sampling the plurality of nodes in a round-robin fashion.

10. The method of claim 9 , wherein the first subset includes a single node of the plurality of nodes.

11. The method of claim 9 , wherein the first subset includes multiple nodes of the plurality of nodes.

12. The method of claim 9 , wherein the first and second subsets are working in parallel on an application.

13. The method of claim 9 , wherein the program code identifies the first and second subsets automatically.

14. The method of claim 9 , wherein the first and second subsets are both sampled at a sampling rate.

15. The method of claim 9 , wherein a sampling rate at which the first subset is sampled is adjustable.

16. The method of claim 9 , wherein the program code determines that the first subset is performing unsatisfactorily.

17. The method of claim 16 , wherein the program code removes the first subset from a subsequent sampling operation in response to determining that the first subset is performing unsatisfactorily.

18. A program product, comprising:

program code configured to sample performance data of a plurality of interconnected nodes comprising a massively parallel computer system, the program code configured to identify a plurality of nodes doing similar work, to distribute an overhead for performance sampling across the plurality of nodes by sampling at different times performance data associated with first and second subsets of the plurality of nodes, and to analyze the performance data, wherein the program code distributes the overhead for performance sampling across the plurality of nodes by sequentially sampling the plurality of nodes in a round-robin fashion; and

a recordable type computer readable medium bearing the program code.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE 1ST ASSIGNEE NAME 50% INTEREST PREVIOUSLY RECORDED AT REEL: 043418 FRAME: 0692. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 1, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.; INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044348/0451 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SERVICENOW, INC.
Reel/Frame 043418/0692 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2007
From: BARSNESS, ERIC LAWRENCE; DARRINGTON, DAVID L.; PETERS, AMANDA E.; SANTOSUOSSO, JOHN MATTHEW
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 018926/0757 →