| Previous | Table of Contents | Next |
The execution continues during the third processor cycle, with the first instruction moving into the execution and effective address calculation stage (stage 3), the second instruction moving into stage 2, and a third instruction entering stage 1. This process continues until after the fifth processor cycle, when the first instruction has completed its execution and has left the pipeline. Thus, any single instruction takes the full five cycles; but after the pipeline is filled, a new instruction completes for every processor cycle. When we say that a processor requires only a single cycle per instruction, we are assuming that the pipeline is full, which clearly is the best case.
During the early 1960s, Cray, who was at Control Data Corporation, was designing the worlds first supercomputer, the CDC 6600. He planned to use pipelining and, for simplicity, he wanted to have all instructions execute in the same amount of time. From our example, we can see that the total instruction time is dictated by the longest running instruction. Instructions that fetch or store operands from memory typically take longer than other types of instructions in a processor. If these memory operations also perform logical or arithmetic operations on the data, the instruction time can become very long.
To keep this total instruction time as short as possible, Cray decided the only memory operations in his design would be to load a register with the contents of a memory location and to store the contents of the registers to a memory location. Any operations on the data would be performed in the registers.
This was a major departure from most other computers that allowed operations on data in memory without using the registers. For example, a System/360 has instructions to allow an operand in memory to be added to another operand in memory and the sum to be put back into memory. This can be a fairly long-running operation, but it is accomplished with a single instruction. This type of instruction is called a memory-to-memory instruction.
Crays machine would take five instructions to accomplish the same operation. First, two load instructions would put the data into two registers. Then an add instruction would add the two register operands and put the sum back into a register. Finally, a store instruction would move the sum from the register to memory. Crays machine took five instructions, but because they could be efficiently pipelined and executed in parallel, the total time required to complete all five was less than an equivalent machine with memory-to-memory instructions required. The disadvantage to Crays machine was the larger number of instructions required to perform the operation.
Introduced in 1964, the CDC 6600 was the first general-purpose load/store machine. Cray understood the interaction between pipelining and instruction set design, and he realized the need to simplify the architecture for the sake of efficient pipelining. RISC processors today use Crays design; this explains why RISC machines, which have only load-and-store memory instructions, are faster than CISC machines, which have a full set of memory instructions. It also explains why programs compiled for a RISC machine are larger.
Crays instruction set design was not the only contribution to high-performance pipelines. Cray introduced hardware in the CDC 6600 to ensure that the pipeline was kept as full as possible. Pipelined machines achieve their maximum performance when the pipeline is full that is, when every stage in the pipeline is executing part of an instruction. In the real world, there are dependencies between instructions in a program. If an instruction in the pipeline uses data stored by another instruction just ahead of it in the pipeline, the data may not be available in time. This causes a stall in the pipeline. Furthermore, all instructions behind the stalled instruction are also stalled, reducing the processors performance.
The CDC 6600 introduced hardware that allowed the processor to look at instructions further back in the instruction stream and to determine whether they could be started before the instruction that had to wait for another result to be stored. This idea of allowing the hardware to re-arrange instructions in the pipeline, called dynamic scheduling, greatly improved the performance of the CDC 6600 by keeping the pipeline as busy as possible.
Another idea used by the supercomputers of the 1960s was branch prediction. A branch instruction can create havoc in a pipeline, causing it to stall until the system can decide the next instruction to use. The idea of branch prediction was to guess, based on experience, where the next instruction would come from when a branch was encountered. Sophisticated hardware was used in the IBM 360/91 to do branch prediction with remarkably good results.
The 360/91 had another interesting hardware characteristic. Bob Tomasulo, an IBM engineer, created a new dynamic scheduling algorithm that was implemented in the hardware. This algorithm was an improvement over the one Cray designed a few years earlier; it eliminated many pipeline stalls by allowing instructions in the processor to execute out of order. An instruction that had to wait for some result no longer stalled instructions behind it in the instruction stream. Tomasulos algorithm required tremendously complex hardware for the time, but it did produce the desired performance improvements.
All of this specialized hardware to optimize the pipeline performance added to the complexity and hardware costs of these systems. This is not a problem for a cost-is-no-object supercomputer, but it did prevent use of these techniques in ordinary systems.
In the late 1960s, John Cocke was working on the design of a fast scientific computer at the IBM Research Laboratory in San Jose, California. Cocke was concerned with the complexity of the hardware needed to do dynamic scheduling to keep the pipelines full. He believed that if much of the responsibility for keeping the pipelines full could be given to the compilers, then the hardware could be greatly simplified. Further, simpler hardware meant lower costs. If it was possible to let the compilers absorb the complexity, then high-performance processing would no longer be the domain of just supercomputers. The idea of RISC was born.
| Previous | Table of Contents | Next |