| Previous | Table of Contents | Next |
Because no design lasts forever, and because a superscalar RISC design with out-of-order and speculative execution is already reaching the point of diminishing returns, we must look at alternatives. One alternative to RISC that is being widely discussed is Very Long Instruction Word (VLIW). The first processor that Intel is developing jointly with HP (Merced) was originally believed to use VLIW technology. Now it appears this processor will use a far more conventional superscalar design. However, it may still incorporate some VLIW concepts. Even if Merced is not a pure VLIW design, lets look at why the VLIW technology has generated so much interest.
VLIW technology has the potential to reverse the trend of superscalar RISC processors toward more on-chip complexity. Earlier, we called these processors Brainiacs because of this added complexity. Simpler designs, such as the Speed Demons, can spin faster and achieve higher MHz ratings. VLIW pushes the complexity back to the compilers, enabling potentially higher-speed processors.
The problem with RISC processors, and the reason for all the complex hardware, is the difficulty of keeping the pipelines full. We saw that most superscalar RISC designs have the capability to dispatch only a small number of instructions, say three or four, in a single cycle, which limits the amount of instruction parallelism we can get in a single processor to the number of instructions that can be dispatched. If I can dispatch only a maximum of four instructions per cycle, the maximum parallelism I can ever achieve is four times and, because of dependencies between instructions as well as branches, the average parallelism is likely to be more like two times or even fewer on some real-world workload. Why cant a superscalar RISC processor dispatch 8 to 16 instructions in a single cycle?
The answer has to do with a RISC processors limitations. First, a RISC processor usually doesnt have enough independent functional units to allow 8 instructions to run in parallel this is a hardware technology limitation. The second limitation is that there is not enough time in a cycle to analyze 8 to 16 instructions, determine which functional units are not busy, and send each instruction to the appropriate functional unit. Taking enough time to do all this would slow down the cycle time and reduce the processor performance. The third limitation is the compilers inability to package 8 to 16 independent instructions that can be dispatched each and every cycle.
Hardware technologies have advanced so that a single-chip processor with 8, 16, or even more functional units is possible. Compiler technology has also improved to the point that more instruction parallelism can be recognized and more independent functional units can be scheduled. But this capability to schedule more instructions is useless if the processor hardware can dispatch only a few instructions per cycle, as a superscalar RISC processor does. To solve this problem, a VLIW design takes the dispatching function away from the processor hardware.
Rather than having the processor hardware analyze each instruction in the instruction stream and then dispatch the instructions one at a time to the functional units, as a RISC processor does, the VLIW compiler creates a separate instruction for each functional unit for each cycle. For example, if we have 16 functional units, the compiler creates 16 instructions for each processor cycle. Each cycle, the VLIW processor fetches from its instruction cache all 16 instructions; but unlike a RISC processor, which has to analyze which functional unit can take which instruction, the VLIW processor simply sends each instruction to its corresponding functional unit. It sends the first instruction to the first functional unit, the second instruction to the second functional unit, and so on. Of course, if the compiler has no work for some functional unit during a given cycle, it still must create a no-operation instruction for that unit. Because the VLIW processor does no thinking, the cycle time of this processor can be shorter than that of an equivalent superscalar RISC processor built from the same hardware technology. This shorter cycle time, plus the increased instruction-level parallelism it can achieve with more usable functional units, gives a VLIW design a performance advantage over the RISC design.
Why is it called a VLIW processor? you might ask. The compiler packages all those independent instructions for each cycle into one very long instruction word hence, the VLIW name. The processor fetches one of these very long words from its instruction cache each cycle. Lets see, if we have 16 instructions, each of which is 4 bytes long, we have a 64-byte- (512-bit) long instruction word. That certainly qualifies as a very long instruction word.
The back end of the compiler (which is similar to the AS/400s translator) for a VLIW processor has the ultimate responsibility to find enough useful work for the processor to do each cycle and to create the instructions to perform that work. If from 4 to 20 useful instructions can be executed per cycle, we can produce throughputs in the giga-instruction-per-second range.
The biggest problem with the VLIW approach is that the back end of the compiler needs to be intimately tied to the hardware. To schedule an instruction for each functional unit in the processor, the compiler needs to know exactly what units are present on the chip and how they relate to one another, which makes it difficult, if not impossible, to use the same compiler output on different chip implementations because there is no binary compatibility.
Early speculation about the approach Intel would use for its Merced chip was to do on-the-fly translation from the x86 and IA-64 instructions to the VLIW instructions. This is similar to the approach Intel uses in its Pentium II and Pentium Pro processors, where the x86 instructions are converted on the fly to a sequence of RISC-like instructions right on the chip. Intel calls these RISC-like instructions micro-ops and describes this technique as dynamic execution. The core of the processor then takes these micro-ops and executes them in a pipelined implementation that looks just like a RISC processor.
Intel was not the first to do this. NexGen, another x86 chip maker now owned by Advanced Micro Devices (AMD), was the first to introduce this idea in its Nx586. NexGen called its internal instructions RISC86 instructions. AMDs K6, another x86-compatible chip, now uses this same design. This dynamic execution approach works for a RISC processor, as these several vendors products demonstrate, but it may not work well for a VLIW processor for reasons related to the amount of parallelism in the instruction stream.
| Previous | Table of Contents | Next |