>linuxsymposium

July 20-23rd, 2005, Ottawa, Canada

Registration

Register/Submit Proposal

Content

Schedule
Presentations
Tutorials
BOFS

Related

Sponsors
Venue
Travel
FAQ

Archives

Proceedings
Photos
2005
2004
2003
2002
2001
2000
1999

Contacts

Information
Home

Building Highly Available Linux Cluster with HA-OSCAR

Ibrahim Haddad (ibrahim@osdl.org)

HA-OSCAR is an Open Source project that aims to provide a combined power of high availability and performance computing. The project’s goal is to improve the existing Beowulf architecture and cluster management systems (e.g. OSCAR, ROCKS, Scyld etc) while providing high-availability and scalability capabilities.

HA-OSCAR enhances a Beowulf cluster system for mission critical grade applications with various HA mechanisms such as component redundancy, self-healing mechanism, failure detection and recovery mechanisms, in addition to supporting automatic failover and fail-back.

Release 1.0 supports new HA capabilities for Linux Beowulf clusters based on the OSCAR 3.0 release from the Open Cluster Group. In this release of HA-OSCAR, we provide an installation wizard GUI and a web-based administration tool, which allows intuitive creation and configuration of a multi-head Beowulf cluster. The new features in the initial release include head node redundancy, self-recovery for hardware, service, and application outages. In addition, we have included a default set of monitoring services to ensure that critical services, hardware components, and important cluster resources are always available. HA-OSCAR also supports new tailored services that can be configured and added via a WebMin-based administration tool.

In this tutorial, we address design and implementation challenges when building HA Beowulf clusters using Linux and Open Source software as the base technology. We will present HA-OSCAR architecture, review the features of the latest release, discuss the implementation of HA and security features, and discuss our experiments covering modeling, and testing performance and availability on real systems. We will also provide a demo given we have enough time.