| Previous | Table of Contents | Next |
IBM first shipped OptiConnect clustering technology to AS/400 customers in mid-1994 to satisfy the growth needs for a very few high-end customers. As I discussed in Chapter 6, the uses for this fiber-optic connection product have expanded tremendously in the past few years. The loosely coupled parallel database support, for example, has enabled some very large installations running applications such as data warehousing and transaction-intensive OLTP environments. Today, there are some huge OptiConnect AS/400 installations around the world. For example, one of the early adopters of the OptiConnect technology has an AS/400 cluster with more than 20,000 attached terminals handling nearly a million transactions per hour. Now thats a big system.
In addition to using an OptiConnect cluster just for growth, most customers with these installations also use the cluster to achieve continuous availability. You can achieve an AS/400 continuous-availability solution with an AS/400 cluster easily enough by adding a mirroring or replication package from a business partner. These software packages from business partners are similar in that they mirror or replicate DB2/400 data and transactions across multiple AS/400s. The AS/400s can be connected via OptiConnect or, for that matter, any other communications link.
In the event of a system or software failure, processing automatically switches to another AS/400 in the cluster. A failover system of this type can result in relatively minor overall performance degradation, and it can minimize user downtime in the event of a failure. These systems also provide 24 × 7 availability for production systems, because backups can be done from the mirrored system in the cluster.
A drawback to an AS/400 cluster is that you cannot manage it as though it were a single system. Clusters from some other vendors do present a single-system image to the programmer and system administrator. A goal for the AS/400e series is to provide this single-system image of the cluster, enabling a user to manage the entire cluster from a single point of control. This and several other enhancements to the e-series cluster support will roll out in stages during Version 4. I should point out that in addition to enhancing continuously available cluster solutions, we will deliver several single-system availability improvements two of which are greatly reduced IPL times and improved save/restore performance as part of Version 4.
I described OptiConnect, one of the new interconnections between systems in a cluster, in Chapter 6. Recall that OptiConnect is implemented as a serial optical SPD bus. For the future we needed a new faster, more fault-tolerant connection. That new connection is the SAN interconnection. In the future, systems in a cluster using SAN will be connected in a fault-tolerant loop configuration. Each system provides two SAN ports that are directly linked to the next system in the loop. A redundant link is used between systems for fault-tolerant operation in case the primary link fails. The SAN protocol provides error detection, packet retry, and alternate-link routing directly in the hardware.
The actual interconnect media for the SAN loop can be either copper or optical fiber, depending on the distance between systems. The optical fiber works for the longer distances. Very high-speed connections for SAN interconnections will be achieved with parallel rather than serial fiber-optics connections. An advanced technology program in Rochester, for example, has demonstrated a parallel fiber-optics connection with 32 fibers running at 500 MHz. That connection provides a throughput of 1 GB per second (assuming half of the fibers are used for the redundant link). That is eight times faster than the 1 Gigabit performance of the fastest OptiConnect implementation. Of course, because SAN is available on only the newest AS/400e systems and servers, OptiConnect will be the only cluster interconnection for a few years.
A likely sequence for the rollout of enhancements for continuous-availability cluster support will be, first, improvements to remote journaling and object replication between systems, to eliminate all single points of failure. These enhancements will help improve the performance and functionality of the mirroring and replication packages from business partners. The next logical step is to allow disks to be switched between systems. Thus, if a primary system fails, the disks that contain the database can be switched to the secondary system. This capability avoids the cost of a duplicate database on the secondary system, but it also removes the capability to do backups from the secondary system. For this reason, some customers may still need mirrored systems. The third step in this rollout is to share a database between systems. The model IBM currently is using internally for database sharing is the System/390 Parallel Sysplex.
The mechanism that enables both disk switching and disk sharing is the Independent ASP (IASP). I described auxiliary storage pools (ASPs) in Chapter 8. Basically, an ASP is a collection of disk devices, where all the disk storage in a pool appears to be a single contiguous area in memory. ASPs contain the various system objects and are used to improve the recovery from a disk failure, which can be isolated to a specific ASP. There is one system ASP and up to 15 user ASPs.
An IASP is a special form of a user ASP. Each IASP is self-contained, so, for example, the system can be IPLed without the IASP being online. Also, an IASP can be taken offline without taking down the entire system. A task in the system gets an exception when accessing an object in an IASP that is offline. IASPs can be attached to a single system or to multiple systems. IASPs attached to multiple systems are treated as a cluster resource and can be switched or shared between systems in the cluster.
The ideas for IASPs came out of some work done in a department I managed shortly after the announcement of the AS/400. Jim Ranweiler at the time was looking for ways to create high-availability AS/400 clusters, and he came up with the IASP concept. We put that work on the shelf until we needed it. Now that continuously available clusters are in our plans, we have been able to dust off that work and put it into the AS/400.
| Previous | Table of Contents | Next |