IP Library Granted Patent US 7,788,540
Granted Patent B2
US 7,788,540 · App. 11/700,992 · Granted Aug 31, 2010

Tracking down elusive intermittent failures

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,788,540
App. No.
11/700,992
Granted
Aug 31, 2010
Kind
B2
Abstract

Computing environments, each executing at least one software program, are monitored for failures occurring during execution of the software program. Information associated with the failure, such as an identification of the software program and a failure type describing the failure, is recorded. The failure information is quantified to report the number of times the software program has failed or the number of times a particular failure has occurred. The quantified data may provide help in prioritizing what program or what failures merit investigation and resolution. Reports may be received from failing computing systems stopped at a state following the occurrence of the failure. In response, hold information is checked to determine whether to instruct the failing computing system to hold a state existing upon the occurrence of the failure.

Claims (77)

1. A computer-implemented method, comprising:

monitoring a plurality of computing environments that are each executing a software program for notification of occurrences of a failure during execution of the software program in each of the computing environments; wherein the failure is an intermittent failure that does not occur each time a same instruction or group of instructions executes within the software program;

recording failure information associated with each of the occurrences of the failure, the failure information including:

identification of the software program; and

a failure type describing the failure; and

quantifying the failure information to maintain a total of at least one of a number of times:

the software program has failed; and

the failure type has been recorded; and

ranking a number of times the software program has experienced the occurrences of failures in each of the computing environments and using the ranking to determine when to address the failure.

2. The method of claim 1 , further comprising:

monitoring execution of a plurality of software programs for occurrences of failures; and

ranking the plurality of software programs according to a number of times each of the plurality of software programs has experienced the occurrences of failures.

3. The method of claim 2 , further comprising providing a graphical view of the ranking of the failures.

4. The method of claim 1 , further comprising:

receiving a report from a failing computing system paused at a failure state following the occurrence of the failure;

checking hold information describing whether to instruct the failing computing system to hold at the failure state; and

instructing the failing system to hold at the failure state when the hold information indicates the failing system is to be held upon the occurrence of the failure reported.

5. The method of claim 4 , wherein the hold information includes one of:

submission data included in initiating the execution of the at least one software program indicating the execution of the at least one software program is to be held at the failure state upon occurrence of the failure; and

when the failure report includes a failure type describing the occurrence of the failure, the hold information includes failure tag data indicating the execution of the at least one software program is to be held upon occurrence of a selected failure type including the failure type in the failure report.

6. The method of claim 5 , further comprising, wherein on at least one second computing environment at least one additional software program executes in cooperation with an additional software program including at least one of another instance of the first software program and a second software program, instructing the second computing system to hold at a current state.

7. The method of claim 4 , further comprising sending a failure message to a user named in the hold information that is to be notified upon the occurrence of the failure.

8. The method of claim 7 , wherein the failure message includes access information configured to facilitate the user being able to access the failing computing system in order to investigate at least one of the failure state and the occurrence of the failure.

9. The method of claim 4 , further comprising instructing the failing computing to hold at the failure state until at least one of:

a time interval has lapsed;

the failure state has been investigated; and

a hold discontinue instruction is given.

10. The method of claim 4 , further comprising:

identifying an original user for whom the failing computing environment had been allocated prior to the occurrence of the failure;

allocating to the original user an additional computing environment to replace the failing computing system being held at the failure state.

11. A computer-readable medium having stored thereon computer-executable instructions, comprising:

monitoring a plurality of computing environments executing at least one software program for notification of occurrences of a failure during execution of the at least one software program in each of the computing environments; wherein the failure is an intermittent failure that does not occur each time a same instruction or group of instructions executes within the at least one software program;

detecting a failure report from a failing computing system paused at a failure state following the occurrence of the failure, the failure report identifying at least one of:

the at least one software program; and

a failure type describing the occurrence of the failure;

checking hold information describing whether to instruct the failing computing system to hold at the failure state; and

instructing the failing system to hold at the failure state for a predetermined time when the hold information indicates the failing system is to be held upon the occurrence of the failure reported.

12. The computer-readable medium of claim 11 , wherein the hold information includes at least one of:

submission data included in initiating the execution of the at least one software program indicating the execution of the at least one software program is to be held at the failure state upon occurrence of a first failure; and

failure tag data indicating the execution of the at least one software program is to be held upon occurrence of a selected failure type indicated in the failure report.

13. The computer-readable medium of claim 11 , further comprising at least one of:

sending a failure message to a user named in the hold information that is to be notified upon the occurrence of the failure; and

providing access information configured to facilitate the user being able to access the failing computing system in order to investigate at least one of the failure state and the occurrence of the failure.

14. The computer-readable medium of claim 11 , further comprising, wherein on at least one second computing environment at least one additional software program executes in cooperation with an additional software program including at least one of another instance of the first software program and a second software program, instructing the second computing system to hold at a current state.

15. The computer-readable medium of claim 11 , further comprising instructing the failing computing to hold at the failure state until at least one of:

a time interval has lapsed;

the failure state has been investigated; and

a hold discontinue instruction is given.

16. The computer-readable medium of claim 11 , further comprising:

identifying an original user for whom the failing computing environment had been allocated prior to the occurrence of the failure;

allocating to the original user an additional computing environment to replace the failing computing system being held at the failure state.

17. The computer-readable medium of claim 11 , further comprising:

recording failure information associated with the occurrence of the failure including at least one of:

the identification of the software program; and

the failure type describing the occurrence of the failure; and

quantifying the failure information to maintain a total of at least one of a number of times:

the software program has failed; and

the failure type has been recorded.

18. A system for facilitating analysis of an occurrence of a failure occurring in a failing software system, comprising:

a plurality of computing environments, each of the plurality of computing environments executing at least one software program being monitored and being configured to generate a failure message reporting an occurrence of failures occurring during execution of the at least one software program in each of the computing environments and provide failure information describing the occurrence of the failure; wherein the failure is an intermittent failure that does not occur each time a same instruction or group of instructions executes within the at least one software program;

a monitoring system in communication with the plurality of computing systems and configured to receive the failure message; and

one of:

record the failure information; and

respond to the failure message by instructing a failing computing system reporting the occurrence of the failure to hold at a failure state existing subsequent to the occurrence of the failure.

19. The system of claim 18 , wherein the system is further configured to at least one of:

quantify the failure information to maintain a total of at least one of a number of times:

the software program has failed; and

the failure type has been recorded; and

rank:

the number of times the software program has failed; and

the number of times the failure type has been recorded.

20. The system of claim 18 , wherein the system is further configured to at least one of:

check hold information describing whether to instruct the failing computing system to hold at the failure state;

instruct the failing system to hold at the failure state when the hold information indicates the failing system is to be held upon the occurrence of the failure reported; and

at least one of:

notify a user named in the hold information that is to be notified upon the occurrence of the failure; and

provide access to the user to the failing computing system to allow the user to investigate at least one of the failure state and the occurrence of the failure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →