Previous Table of Contents Next


This memory bus contention shows up as reduced performance for each additional processor in the SMP configuration. For example, adding a second processor does not increase the total performance by 100 percent; the actual increase will be somewhat less because of the memory contention. Adding a third and fourth processor reduces the per processor performance even further. By the time you get to eight or more processors, adding another processor may not give any more performance from the SMP configuration.

Memory Organizations

Before we go further, a brief explanation of memory organizations for multiprocessor systems is in order. Three organizations are of interest to us: centralized shared memory, distributed memory, and distributed shared memory.

A centralized shared-memory machine has a central memory with multiple processors all sharing this memory. Typically, up to a few dozen processors are connected to the shared memory with a single bus. This organization is more commonly called SMP, and it is exactly the one we have described in the preceding paragraphs. The advantages of an SMP configuration are that all processors share the same memory, and the amount of time required for any processor to get to any part of the central memory is the same. Because of this latter advantage, we often call SMP configurations uniform memory access (UMA) machines. The disadvantage of SMP is that a limited number of processors can be supported.

A distributed memory machine has the memory distributed among a number of nodes, where a node contains a small number of processors connected to the node’s memory in an SMP fashion. An example is the AS/400 OptiConnect cluster that we discuss in Chapter 6. A distributed memory machine is sometimes called a shared-nothing design, because memory is not shared among the nodes. Communication between nodes is accomplished by passing messages back and forth. The advantage of a distributed memory is that it can get very big, as long as applications do not have to share memory.

Further, if each processor in a distributed-memory machine can be instructed to perform the same operation or the same program on multiple sets of independent data, the configuration is called a massively parallel processor (MPP). An example of an MPP system that uses one processor per node is the IBM SP2. The SP2 also has a very nice message-passing switch to allow processors to communicate with one another very rapidly. MPP systems can become enormous with many thousands of processors. The disadvantage is that MPP is only useful for certain types of applications, such as parallel database operations or scientific computing applications where sharing of data is not required.

A third configuration, called a distributed shared-memory machine is a variant of the distributed memory machine. Here, a common address space is shared among all the nodes. Again, the nodes have one or more processors in an SMP arrangement. The difference between this machine and the distributed memory machine is that any processor can get at any part of the memory with the distributed shared-memory machine. The amount of time it takes to get to some part of the memory, however, varies depending upon where that part of memory is physically located in the cluster. For this reason, these machines are called non-uniform memory access (NUMA). We study NUMA in Chapter 12.

I have introduced these memory configurations to help explain the new memory subsystem being used in today’s AS/400e series. This AS/400 memory subsystem is designed for an SMP configuration, because this is the best fit for a commercial workload where the processors need to share memory; but as we will see, other memory configurations also can use this memory subsystem.

We’d Rather Switch than Fight

As a part of the 10-year long Accelerated Strategic Computing Initiative (ASCI) started in 1995, the U. S. Department of Energy (DOE) requested computer vendors to propose designs for some of the world’s most powerful computers. The goal of ASCI is to develop tera-scale computers that can initially be used to simulate nuclear devices without resorting to nuclear testing. Tera-scale computing, which is the official name for a trillion calculations per second, is expected to have many commercial and scientific applications in the next century. These computer designs are being deployed for use and evaluation at the three DOE national laboratories associated with the ASCI project.

The first stage of the ASCI project, called ASCI Option Red, was for a large MPP configuration of processors arranged in the traditional distributed-memory model we just described. Intel proposed and was awarded a contract for a computer with 9,072 Pentium Pro processors, 283 GB of memory, and 2 TB of disk storage. This design is a shared-nothing MPP design, and it was installed at the Sandia National Laboratory in New Mexico. The goal was to wring one trillion floating-point operations per second (a teraflop) out of this $55M one-of-a-kind computer. In December 1996, the Intel DOE computer achieved its goal.

The DOE also wanted to address the limitations of the two common multiple-processor architectures (SMP and MPP). As we have seen, SMP designs based on buses do not scale much beyond 32 processors, but they work very well for most applications. MPP designs are more finicky to program and fit only certain applications. They also become clogged and slowed if they need to access data spread all over the system. Thus, the DOE requested proposals for a new scalable SMP design that it called ASCI Option Blue.

Two of the companies that made proposals were awarded contracts to deliver their systems late in 1998: IBM and Cray Research, which has been acquired by Silicon Graphics Incorporated (SGI). The proposed designs from these companies each showed promise, so the DOE awarded contracts to have both systems built. IBM will install its design, called ASCI Blue Pacific, at the Lawrence Livermore National Laboratory in California. SGI/Cray will install its design, called the ASCI Blue Mountain project, at Los Alamos National Laboratory in New Mexico. The performance goal for both of the Option Blue computers is greater than 3 teraflops.

The IBM design uses compact SMP nodes with eight processors; these nodes are then connected using SP2 message-passing switches. The SGI/Cray design is more complex, with a combination of interconnections and operating-system technologies designed to give the image of a single, SMP-like system, even though data is physically distributed around the system; this will be a NUMA design.


Previous Table of Contents Next

Copyright © NEWS/400 Books