| Previous | Table of Contents | Next |
The six chips that make up a single processor include a Processing Unit (PU) chip, a Floating-Point Unit (FPU) chip, and four copies of a Main Store Control Unit (MSCU) chip. The PU chip contains the instruction cache, the branch unit, and the fixed-point unit. The instruction cache has 8K, organized 32 bytes wide. The cache can fetch 32 bytes (eight instructions7) per cycle from memory. The 32-byte data paths are designed to transfer large amounts of data in a single cycle. Even the register stack in the fixed-point unit is designed to load or store four 64-bit registers in a single cycle.
7All PowerPC instructions are 32 bits wide for both the 32- and 64-bit processors.
The FPU is contained on a single chip and supports the IEEE standard for floating-point. The design of the FPU is such that it can produce a result every cycle, giving this processor its very high floating-point performance. Four instructions per cycle are passed from the PU chip to the FPU chip on the 16-byte P-bus. All store data from the PU is also passed across the P-bus on its way to the data cache. Data for the FPU is fetched from the data cache at a rate of 32 bytes per cycle. All store data from either the FPU or the PU is passed to the data cache across the 16-byte store bus.
The MSCU provides the data cache as well as the interface to the memory. All four chips work together to provide the 256K data cache and the interfaces to the data buses shown in Figure 2.4. The data cache accesses are pipelined so that 32 bytes of data are fetched and 16 bytes of data are stored each cycle. The MSCU supports multiprocessing by providing cache coherency across multiple processors.
In total, there are five pipelines in the PU and FPU, but only four instructions can be dispatched each cycle. The four instructions are
The floating-point instructions are executed in the FPU, but they cannot be dispatched at the same time as a fixed-point logical, shift, or rotate instruction, which executes in the PU. Note that the load/store pipeline does both fixed and floating-point loads and stores. The multiple pipeline implementation allows parts of several instructions to be executing simultaneously.
Muskie also supports tightly coupled, shared memory, so it can support symmetric multiprocessing (SMP) configurations. Up to four-way multiprocessing was supported for this first-generation processor.
Besides being a fast RISC processor, Muskie was optimized for the needs of a commercial processor. A few of these characteristics will illustrate the point:
| Previous | Table of Contents | Next |