| Previous | Table of Contents | Next |
For the first time, the AS/400e series offers SMP configurations in the middle of the product line. The Millennium box can support up to 4-way SMP configurations using Apache processors. The high-end Mako box supports up to 12-way SMP configurations. The former high-end package (Key Largo) was designed to accommodate up to four Muskie processors. Mako is a much taller box (at 1.5 meters, it is reminiscent of the original AS/400 white racks) that is designed to hold more processors, main memory, and disk drives than the smaller boxes. Expansion units are used with both Millennium and Mako boxes for even larger configurations. The new packages, which are intended to carry the AS/400e series to the year 2000, were designed jointly with the RS/6000 division so that both systems can share more components.

Earlier, we looked at the Muskie processor and saw how it was designed to support a data-intensive commercial processing environment. A 4-way superscalar design and the 16-byte- (128 bit-) and 32-byte- (256 bit-) wide buses clearly indicated that Muskie was designed to handle massive amounts of data and instructions. Imagine for a moment that the entire Muskie processor along with a couple of 64K caches for instructions and data could be packaged on a single CMOS chip. Add to that Apaches full PowerPC architecture and the capability to handle 12-way or even 16-way SMP. What would you have? You would have our fourth-generation Northstar processor, of course.
We will hold off on our discussion of futures until Chapter 12. There, we look at several technologies, including processor technologies that IBM is investigating for the AS/400. For now, we can be certain that new processors over the next few years will be single-chip CMOS designs that can be packaged in the newest black boxes. We also know they will take advantage of the new memory subsystem announced with Apache.

One of the biggest problems any processor has is keeping busy. Over the past several years, processor performance has continued to grow dramatically, often doubling every two years. Memory and I/O performance have not kept pace. In 1991, I bought a new IBM PC for my own use at home. It had a 20 MHz Intel 386 processor, 70-nanosecond memory, and a hard-disk drive with a 16-millisecond access time. The little 386 processor didnt suit my needs for very long; so like many other people, I began the endless trek for the latest PC hardware, acquiring systems with 486 and Pentium processors. My latest purchase, which is already old technology, has 200 MHz Pentium Pro processors, 60-nanosecond memory, and disk drives with 8.5-millisecond access time. Look at what has happened over the past few years the differential between the speed of the processor and the speed of the memory-I/O has increased dramatically. The possibility exists that the memory and I/O may now have difficulty feeding enough information to the processor to keep it busy, and the processor may spend many of its high-speed cycles waiting for instructions and data to be fetched.
The universal technique used to compensate for the performance differential between processors and main memories is to use cache memories. A cache memory, as we have seen, is a high-speed memory used to store blocks of instructions and data that the processor has recently accessed. The assumption is that the processor will access these same blocks again in the near future. Caches work because most programs exhibit what we call spatial and temporal locality. This simply means that it is highly likely that a program will access an instruction or a piece of data in memory that is in close proximity to a previous access (spatial locality) and that the program will make that access in a short period of time (temporal locality). As you might expect, some programs exhibit higher levels of locality than others. A program that does lots of looping or reprocesses the same data has a very high level of locality. Commercial applications do little of this type of processing, and they exhibit somewhat lower levels of locality. Large caches and hierarchies of caches can help keep instructions and data available for the processor to use even with low levels of locality.
With a cache hierarchy, if the information the processor wants is not in the L1 cache, the processor looks in the L2 cache. If the information still isnt found there, the processor goes to the main memory. Throughout this process of looking for the piece of information and fetching it into a processor register, the processor is idle. We usually talk about these idle times as stall cycles in the processor. In the case in which the information is not even in the memory and the system has to go out to the disk, the processor is directed to work on another job in the system rather than wait. This case is usually called a page fault, which we cover in Chapter 8.
Muskie achieved a high level of performance by using a small 8K L1 instruction cache and a 256K L1 data cache. This large data cache for Muskie is implemented with four separate high-speed 64K memory chips. Because they are single-chip processors, Cobra and Apache have much smaller on-chip L1 caches (4K and 8K for instructions and data, respectively). These on-chip caches are backed up with larger L2 caches implemented on separate chips. Apache achieves some of its performance through the use of very large L2 caches of 4 to 8 MB.
Even with large caches and hierarchies of caches, most processors still have lots of stall cycles waiting for memory. The problem has to do with the bus between the processors and the main memory. Most systems, including the early AS/400s, use a single memory bus between the processors and the main memory. This is usually a fairly high-speed bus, but often it runs slower than the processor itself. As an example, my 200MHz Intel Pentium Pro has a 66 MHz memory bus. This means when the processor goes to memory it stalls three cycles for every bus cycle needed to complete the memory operation. It is easy to see these stall cycles can add up.
An SMP configuration can greatly aggravate the situation. Multiple processors now are trying to use the single memory bus and the main memory. Because only one processor at a time can use the bus, all other processors that also want to use it must wait. If the I/O uses this same memory bus, the situation gets even worse.
| Previous | Table of Contents | Next |