Sun Microsystems, Inc.
spacerspacer
spacer www.sun.com docs.sun.com |
spacer
black dot
 
 
  Previous   Contents   Next 
   
 
Chapter 4

Frequently Asked Questions

This chapter includes answers to the most frequently asked questions about the SunPlex system. The questions are organized by topic.

High Availability FAQs

  • What exactly is a highly available system?

    The SunPlex system defines high availability (HA) as the ability of a cluster to keep an application up and running, even though a failure has occurred that would normally make a server system unavailable.

  • What is the process by which the cluster provides high availability?

    Through a process known as failover, the cluster framework provides a highly available environment. Failover is a series of steps performed by the cluster to migrate data service resources from a failing node to another operational node in the cluster.

  • What is the difference between a failover and scalable data service?

    There are two types of highly available data services, failover and scalable.

    A failover data service runs an application on only one primary node in the cluster at a time. Other nodes might run other applications, but each application runs on only a single node. If a primary node fails, the applications running on the failed node fail over to another node and continue running.

    A scalable service spreads an application across multiple nodes to create a single, logical service. Scalable services leverage the number of nodes and processors in the entire cluster on which they run.

    For each application, one node hosts the physical interface to the cluster. This node is called a Global Interface (GIF) Node. There can be multiple GIF nodes in the cluster. Each GIF node hosts one or more logical interfaces that can be used by scalable services. These logical interfaces are called global interfaces. One GIF node hosts a global interface for all requests for a particular application and dispatches them to multiple nodes on which the application server is running. If the GIF node fails, the global interface fails over to a surviving node.

    If any of the nodes on which the application is running fails, the application continues to run on the other nodes with some performance degradation until the failed node returns to the cluster.

File Systems FAQs

  • Can I run one or more of the cluster nodes as highly available NFS server(s) with other cluster nodes as clients?

    No, do not do a loopback mount.

  • Can I use a cluster file system for applications that are not under Resource Group Manager control?

    Yes. However, without RGM control, the applications need to be restarted manually after the failure of the node on which they are running.

  • Must all cluster file systems have a mount point under the /global directory?

    No. However, placing cluster file systems under the same mount point, such as /global, enables better organization and management of these file systems.

  • What are the differences between using the cluster file system and exporting NFS file systems?

    There are several differences:

    1. The cluster file system supports global devices. NFS does not support remote access to devices.

    2. The cluster file system has a global namespace. Only one mount command is required. With NFS, you must mount the file system on each node.

    3. The cluster file system caches files in more cases than does NFS. For example, when a file is being accessed from multiple nodes for read, write, file locks, async I/O.

    4. The cluster file system supports seamless failover if one server fails. NFS supports multiple servers, but failover is only possible for read-only file systems.

    5. The cluster file system is built to exploit future fast cluster interconnects that provide remote DMA and zero-copy functions.

    6. If you change the attributes on a file (using chmod(1M), for example) in a cluster file system, the change is reflected immediately on all nodes. With an exported NFS file system, this can take much longer.

  • The file system /global/.devices/node@<nodeID> appears on my cluster nodes. Can I use this file system to store data that I want to be highly available and global?

    These file systems store the global device namespace. They are not intended for general use. While they are global, they are never accessed in a global manner--each node only accesses its own global device namespace. If a node is down, other nodes cannot access this namespace for the node that is down. These file systems are not highly available. They should not be used to store data that needs to be globally accessible or highly available.

Volume Management FAQs

  • Do I need to mirror all disk devices?

    For a disk device to be considered highly available, it must be mirrored, or use RAID-5 hardware. All data services should use either highly available disk devices, or cluster file systems mounted on highly available disk devices. Such configurations can tolerate single disk failures.

  • Can I use one volume manager for the local disks (boot disk) and a different volume manager for the multihost disks?

    This configuration is supported with the Solaris Volume Manager software managing the local disks and VERITAS Volume Manager managing the multihost disks. No other combination is supported.

Data Services FAQs

  • What SunPlex data services are available?

    The list of supported data services is included in the Sun Cluster 3.1 Release Notes.

  • What application versions are supported by SunPlex data services?

    The list of supported application versions is included in the Sun Cluster 3.1 Release Notes.

  • Can I write my own data service?

    Yes. See the Sun Cluster 3.1 Data Services Developer's Guide and the Data Service Enabling Technologies documentation provided with the Data Service Development Library API for more information.

  • When creating network resources, should I specify numeric IP addresses or hostnames?

    The preferred method for specifying network resources is to use the UNIX hostname rather than the numeric IP address.

  • When creating network resources, what is the difference between using a logical hostname (a LogicalHostname resource) or a shared address (a SharedAddress resource)?

    Except in the case of Sun Cluster HA for NFS, wherever the documentation calls for the use of a LogicalHostname resource in a Failover mode resource group, a SharedAddress resource or LogicalHostname resource may be used interchangeably. The use of a SharedAddress resource incurs some additional overhead because the cluster networking software is configured for a SharedAddress but not for a LogicalHostname.

    The advantage to using a SharedAddress is the case where you are configuring both scalable and failover data services, and want clients to be able to access both services using the same hostname. In this case, the SharedAddress resource(s) along with the failover application resource are contained in one resource group, while the scalable service resource is contained in a separate resource group and configured to use the SharedAddress. Both the scalable and failover services may then use the same set of hostnames/addresses which are configured in the SharedAddress resource.

 
 
 
  Previous   Contents   Next