| Previous | Table of Contents | Next |
We have already spent enough time on the centralized shared-memory model of the AS/400. The UMA memory subsystem with the crossbar memory switches I discussed in Chapter 2 and variations of that subsystem will easily enable 16-way SMP configurations with the very high-performance processors planned for the AS/400e series. Beyond Version 4, we may even push the SMP configurations to 20- or 24-way.
For very large configurations, we will rely on clusters of SMP nodes. In Chapter 11, I laid out the sequence of AS/400 cluster implementations, from shared-nothing where each system has its own disk drives, to switching disks between the systems, to sharing disks among all systems in the cluster. Once we have the capability to share disks in a single cluster pool of disk drives, as I described through the use of independent ASPs, we can consider sharing the memory across the nodes. Our first NUMA machine would do this.
The interest in NUMA for the AS/400 goes back several years. Dick Booth, a Rochester engineer, had taken an assignment to work on multiprocessors at IBM Research in the early 1990s. While he was there, an idea for a NUMA system design emerged. Dick originally called the design firmly coupled multiprocessing, because it fit somewhere between loosely coupled (which we now call MPP) multiprocessing and tightly coupled (which we now call SMP) multiprocessing today we simply refer to this design as NUMA. Dick believed the idea would work for the AS/400 and brought it back to Rochester.
Dick convinced others in Rochester that this was a good idea, and together they began to sell the idea. They established a joint project with IBM Research in 1991 and started work on a prototype system. Like most new ideas, this one met with some skepticism. The group persisted, finishing the prototype and demonstrating that the idea did, indeed, work for an AS/400. Slowly, the technical community began to come around. Today, members of that same group are working on the NUMA implementations for tomorrows AS/400.
At least two NUMA implementations are possible for the AS/400. The first of the two is called a cache-coherent non-uniform memory access (CC-NUMA), and the second is called a cache-only memory architecture (COMA). The specific implementation details and performance considerations for these designs are the topic of many papers in the computer-architecture arena. Several universities and research laboratories have been investigating variations on each of the designs since the early 1990s. Some computer companies, such as Silicon Graphics, Inc. (SGI), Sequent, and Convex, are already shipping highly scalable CC-NUMA servers.
Without getting into too much technical detail about the implementations of these two designs, I want to introduce them briefly and give some background on the designs to show what you might expect in a future AS/400 configuration. Both configurations use a directory-based cache-coherence protocol to support the shared-memory view, even though the main memory is distributed across the nodes. Simply stated, each node contains a directory showing where all the pages in the globally addressable main memory (local or remote) can be found.
This is different from the snoopy bus-based coherence used for the L2 cache memories in an SMP node, as discussed in Chapter 2. The same data for a shared memory page may at any time be contained simultaneously in the caches of two or more processors in the SMP node. When one processor changes that data in its cache, some way must be provided to update the copies in the caches of other processors. Cache coherence means all copies are up-to-date. With the cache-snooping protocol, the cache directory for each processor keeps track of only the pages in its own cache. Whenever the processor makes a change to the data in its cache, this change is broadcast across the snoopy bus to all other processor caches, so they can be updated if they contain shared data. In this way, each cache can snoop on all changes to the other caches, and cache coherency is maintained. Keeping multiple copies of data in processor caches does ensure the same access time for all data (which is why we call this a UMA design), but the need to broadcast all changes becomes a bottleneck when too many caches are linked together.
A directory-based cache-coherence protocol eliminates the need to broadcast changes, because the directories contain information only about which node contains the data, and not about the data itself. Each node keeps track only of its own local data; there is no duplication of shared data. Referencing data contained in the local node is faster than referencing data on a remote node; hence, the name NUMA. Remember that nodes share the address space, not the memories. Part of the total address space is assigned to each node in the cluster. Extending the address space and updating the directories enables the addition of more nodes to the cluster.
Also remember that the directories we are discussing contain only information about data in the memory for the various nodes. Data on disk is outside of this structure. Simply put, the hardware keeps track of what is in the memories, while the operating system is responsible for keeping track of what is on the disks.
The early implementations of NUMA with the directory-based cache coherence had a relatively large ratio of remote-to-local cache miss times. When a processor in the node encounters an L2 cache miss, the amount of time required to get the data from the memory of a remote node may be far greater than the time to get the data from the memory of the local node containing the processor. For example, an early Sequent machines remote misses were 10 times slower than local misses. To get reasonable performance from such a design, the application data must be carefully distributed across the nodes to keep the number of remote accesses to an acceptably low level. The usual target is to have no more than 10 percent remote accesses. Thus, scaling from an SMP implementation to a distributed cluster implementation may require changes to the application programs and redistribution of the application data.
| Previous | Table of Contents | Next |