| Previous | Table of Contents | Next |
Looking closer, we see that the IBM ASCI Blue Pacific computer will have 512 8-way SMP nodes. The 4,096 processors used in the design will be very high-performance, 64-bit PowerPC designs. The actual processor to be used is the IBM Austin version of Belatrix called the 630. This processor has the high NIC performance that exactly fits the type of problems anticipated for the DOE computer.
Communications between nodes in the ASCI Blue Pacific computer will be accomplished with a new higher-speed SP2-style message-passing switch. The memory subsystem that allows the processors within a node to share memory very efficiently will use a new 128-bit memory crossbar switch.9 A memory subsystem built using these new switches allows multiple processors to concurrently access the memory in the node. This memory subsystem provides a UMA configuration that eliminates the bottleneck of the memory bus used in most SMP configurations.
9Crossbar switches have been used for decades within telephone switching exchanges to connect a set of incoming lines to a set of outgoing lines in an arbitrary way. Any incoming line can be connected to any non-busy outgoing line at any time to provide simultaneous connections.
I have described the DOE project to introduce the new memory subsystem used for the ASCI Blue Pacific SMP nodes. The first UMA memory subsystem to use the 128-bit memory crossbar switches was designed in Rochester, and the SMP Apache designs are currently using this arrangement. Instead of using a single bus between the memory and the processors L2 cache, as previous AS/400 SMP configurations have, Apache uses the memory crossbar switches. With the capability to have multiple concurrent memory accesses per cycle, massive amounts of data can be moved in and out of the shared memory to keep the processors in a large SMP configuration busy.
Figure 2.6 shows a 12-way configuration. Four Apache processors together with four L2 caches are packaged on a single card. A 12-way system uses three of these cards. The L2 caches on the cards each contain either 4 MB or 8 MB and have a cycle time of 8 nanoseconds. Thus, in a single processor cycle, 16 bytes can be transferred between the L2 cache and either the L1 data cache or the L1 instruction cache on the Apache chip (see Figure 2.5).
Figure 2.6 12-Way SMP Configuration
The main memory for this configuration can contain up to 20 GB. Each memory card can contain up to 1 GB, so 20 memory cards are shown in Figure 2.6. Notice in the configuration that there are four memory banks, each with an even number of memory cards to enable memory interleaving, a technique in which consecutive blocks of data in memory can be simultaneously accessed across all memory banks. For example, with the 8-byte interface to each memory card, 32 consecutive bytes can be read or written to the four memory banks at the same time (bytes 0 7 are in bank 1, bytes 8 15 are in bank 2, and so forth).
The four memory crossbar switches shown in the UMA memory subsystem handle the connections between the L2 caches and the main memory cards. Three 6xx data buses, one from each processor card, connect the 12 processors to each of the four memory switches. These 128-bit data buses have a cycle time of 12 nanoseconds (1.5 times the processor cycle time). An additional 6xx data bus connects the I/O to each of the memory switches. Each switch has two independent 128-bit interfaces to the memory cards.
With the configuration shown in Figure 2.6, multiple concurrent accesses to main memory can occur on every memory cycle. Such a structure essentially eliminates the bottleneck a single memory bus causes in a traditional SMP system. Note this is only one configuration for processor cards, switches, and memory cards. Different models within the AS/400 product line will use different combinations of these components. For example, a 4-way SMP configuration might use one processor card, two switches, and four memory banks.
Notice also that only the data lines use the switches. The address lines to all L2 caches (shown in the figure as the 6xx address lines) are shared, allowing normal cache snooping, a technique in which each L2 cache controller continuously monitors all addresses on the shared address bus. In addition to monitoring the addresses (called snooping), the controllers check to see whether the address on the bus is in their cache. If so, the corresponding data in their cache is invalidated. In this way, coherency of information in all caches can be assured because only one processor at a time can change shared data.
Figure 2.6 also shows the common I/O subsystem. This design uses a System Area Network (SAN) to attach an I/O expansion unit to the Mako box that houses this 12-way configuration (two of these interfaces are shown in the figure). In Chapter 10, we will see how SANs are used to support the different I/O interfaces to the AS/400e series and to connect one system to another.
To recap, the reason for using the crossbar switches is to improve SMP efficiency. We often talk about SMP efficiency, by which we mean the percentage of performance improvement we get when we add another processor to an SMP configuration. Many shared memory systems see efficiency in the range of 70 percent for four to eight processors. With the new memory subsystem, the Apache processors should see efficiencies in the range of 85 to 90 percent.
| Previous | Table of Contents | Next |