IP Library Granted Patent US 8,578,214
Granted Patent B2
US 8,578,214 · App. 13/112,775 · Granted Nov 5, 2013

Error handling in a virtualized operating system

Inventors: Laurent Dufour (Plaisance du Touch, FR); Khalid Filali-Adib (Austin, TX); Perinkulam I. Ganesh (Round Rock, TX); Balamurugan Ramajeyam (Chennai, IN); Kavitha Ramalingam (Bangalore, IN); David W. Sheffield (Austin, TX)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,578,214
App. No.
13/112,775
Granted
Nov 5, 2013
Kind
B2
Abstract

When moving workload partitions (WPARs) from machine to machine, operating systems may encounter errors that prevent successful WPAR migration. Recording and reporting errors can be challenging. To move WPARs, the operating system may employ a plurality of software components, such as code residing in user space (e.g., application programs, OS libraries, and shell scripts), code residing in the operating system's kernel, and code residing on remote machines. Embodiments of the invention include a framework that enables all the software components to record errors. The framework can also report the errors to users and processes.

Claims (56)

1. A method for logging errors that arise while moving a workload partition of an operating system from a source machine to a destination machine, the method comprising:

halting processes executing in the workload partition;

determining state information of the processes, wherein the determining is performed by a first group of one or more modules residing on the source machine in memory space assigned to a kernel of the operating system;

detecting, based on the state information, a first error affecting movement of the workload partition from the source machine to the destination machine, wherein the detecting is performed by one or more of the first group of modules;

writing a first message into a log buffer stored on the source machine in the memory space assigned to the kernel, wherein the first message describes the first error, and wherein the writing the first message occurs via a call, by one or more of the first group of modules, to an error logging framework residing in the memory space assigned to the kernel;

detecting a second error affecting the movement of the workload partition from the source machine to the destination machine, wherein the detecting the second error is performed by a second group of one or more modules residing on the source machine in memory space assigned to user programs; and

writing a second message into the log buffer, wherein the second message describes the second error, and wherein the writing of the second message occurring via a call by the second group of modules to the error logging framework.

2. The method of claim 1 further comprising:

reading the first and second messages from the log buffer; and

presenting the first and second messages on an output device.

3. The method of claim 2 further comprising:

translating the first and second messages into a selected language.

4. The method of claim 1 further comprising:

determining a list of processes residing on the source machine to notify about the movement of the workload partition from the source machine to the destination machine;

notifying the processes of the list that the workload partition is moving.

5. The method of claim 1 , wherein the second group of modules include one or more of application programs and shell scripts.

6. The method of claim 1 , wherein the first and second errors will cause movement of the workload to fail.

7. An computer comprising:

a processor;

a memory configured to include a user space and a kernel space, wherein the user space includes a workload partition including processes;

a first group of modules residing in the user space, the first group of modules configured to

migrate the workload partition to a destination computer; and

detect a first error affecting the migration of the workload partition to the destination machine; and

an error logging framework residing in the kernel space of the memory, the error logging framework configured to

write a first message into a log buffer residing in the kernel space of the memory, wherein the first message describes the first error, and wherein writing the first message occurring via a call, by the first group of modules, to the error logging framework.

8. The computer of claim 7 comprising:

a second group of modules residing in kernel space of the memory, the second group of modules configured to

detect a second error affecting the migration of the workload partition the destination computer; and

writing a second message into the log buffer, wherein the second message describes the second error, and wherein writing of the second message occurring via a call by the second group of modules to the error logging framework.

9. The computer of claim 8 , wherein the second group of modules is further configured to

determine a list of processes residing on the source machine to notify about the migration of the workload partition from the source machine to the destination machine; and

notify the processes of the list that the workload partition is moving.

10. The computer of claim 7 , wherein the error logging frame work is further configured to

read the first and second messages from the log buffer; and

present the first and second messages on an output device of the computer.

11. The computer of claim 7 , wherein the error logging frame work is further configured to translate the first and second messages into a selected language.

12. The computer of claim 7 , wherein the first group of modules include one or more of application programs and shell scripts.

13. The computer of claim 7 , wherein the first and second errors will cause migration of the workload partition to fail.

14. A computer program product for logging errors that arise while moving a workload partition of an operating system from a source machine to a destination machine, the computer program product comprising:

a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code configured to, halt processes executing in the workload partition;

determine state information of the processes, wherein the determination is performed by a first group of one or more modules residing on the source machine in memory space assigned to a kernel of the operating system;

detect, by the first group of modules, a first error affecting the movement of the workload partition from the source machine to the destination machine, wherein the detection of the first error results from the determining the state information; and

record a first message into a log buffer stored on the source machine in the memory space assigned to the kernel, wherein the first message describes the first error, and wherein recordation of the first message occurring via a call, by the first group of modules, to an error logging framework residing in the memory space assigned to the kernel.

15. The computer program product of claim 14 , wherein the computer readable program code is further configured to

detect a second error affecting the movement of the workload partition from the source machine to the destination machine, wherein the detecting is performed by a second group of one or more modules residing on the source machine in memory space assigned to user programs; and

write a second message into the log buffer, wherein the second message describes the second error, and wherein the writing of the second message occurring via a call by the first group of modules to the error logging framework.

16. The computer program product of claim 15 , wherein the computer readable program code is further configured to:

read the first and second messages from the log buffer; and

present the first and second messages on an output device.

17. The computer program product of claim 15 , wherein the second group of modules include one or more of application programs and shell scripts.

18. The computer program product of claim 14 , wherein the computer readable program code is further configured to:

translating the first message into a selected language.

19. The computer program product of claim 14 , wherein the computer readable program code is further configured to:

determine a list of processes residing on the source machine to notify about the movement of the workload partition from the source machine to the destination machine;

notify the processes of the list that the workload partition is moving.

20. The computer program product of claim 14 , wherein the first error will cause movement of the workload to fail.

Assignments (2)
CONVEYOR IS ASSIGNING UNDIVIDED 50% INTEREST Recorded Jan 11, 2018
From: INTERNATIONAL BUSINESS MACHINES
To: SERVICENOW, INC.; INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045047/0229 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2011
From: DUFOUR, LAURENT; FILALI-ADIB, KHALID; GANESH, PERINKULAM I.; RAMAJEYAM, BALAMURUGAN; RAMALINGAM, KAVITHA; SHEFFIELD, DAVID W.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 026392/0943 →
Priority Claims (1)
EP 10305971 · Sep 9, 2010 · regional
Continuity (1)
Related Publication 20120066556A1 · Mar 15, 2012