| Previous | Table of Contents | Next |
At about the same time, a group of researchers at Yale University proposed creating a machine with a very wide instruction (512 bits), and they named this type of machine a VLIW. A machine from Multiflow Computers commercialized the ideas developed at Yale, but Multiflow Computers eventually failed because of a lack of financial backing. In 1993, HP licensed the patent portfolio from the defunct computer company.
Rochester continued to have an interest in VLIW machines in the late 1980s, especially as this technology related to the HMC. After IBM announced the AS/400, we started work on the processor design for the next generation of systems. VLIW was a part of that design.
Dave Luick was one of the leaders of that VLIW effort. Dave, one of our original processor designers, first led the design of the System/38 Model 7 processor, and he has been involved with every processor since that time. Always one to push the limits of high-performance processor technology, he was very interested in how to apply some of the VLIW techniques to the HMC. The C-RISC processor, which I discussed in Chapter 2, was being designed as the processor for the Advanced Series before we adopted the PowerPC technology. C-RISC had some RISC capabilities, but it also had an HMC with many characteristics of a VLIW machine, thanks to Dave and others who shared the same vision for processor design.
In 1991, Dave was a member of our 10-person team that was to evaluate the use of PowerPC processors for the AS/400. After we decided to go with the PowerPC technology, he and others turned their attention to building a PowerPC-compatible VLIW machine. Because VLIW is so very dependent on compiler technology, this team immediately formed a joint effort with IBM Research. In Research, a group of people was working on VLIW, but this group needed a platform that could use this technology. Most systems have great difficulty adopting such a new technology without creating a negative impact on their customer base. But the AS/400s technology independence makes that a non-issue. We could introduce VLIW into the AS/400 with no customer disruption.
The VLIW work in Rochester showed the tremendous potential this technology had to increase the AS/400 processor performance. First, the processor cycle times can be decreased because of the simpler structure the design is more like that of a Speed Demon; for a given technology, processor chips that run at twice the speed of a standard PowerPC are possible. Second, it is possible to achieve far more parallelism than with a superscalar RISC design; instead of only 5 or 6 pipelines, 16 or even more on a single chip is an achievable goal in the next couple of years.
The VLIW work is currently on hold in Rochester for a couple of reasons. First, we agreed to use a common processor technology in both the AS/400e series and RS/6000 product lines. Although the AS/400, because of its technology independence, can incorporate a radically new technology such as VLIW, other systems such as the RS/6000 cannot. But both systems can use PowerPC RISC processors.
We did for a while consider building a PowerPC processor with a VLIW core. Both the AS/400 and the RS/6000 could use such a design. This design would also let us build a new AS/400 translator that could either target the PowerPC interface of the processor or bypass the translation step and target the VLIW core directly. Over time, we would rewrite the various components in SLIC to run natively in the VLIW core. Until then, they would run using the PowerPC interface. Likewise, observable application programs could be retranslated into VLIW; nonobservable programs could still run as PowerPC programs.
Because of the controversy surrounding the efficiency of the translation process to a VLIW core, we have suspended the work on a PowerPC processor with a VLIW core. We will have to wait to see how well the Intel Merced incorporates VLIW technology. Some of us have even suggested that we might want to consider porting the AS/400 to this new 64-bit Intel processor. Now that would be an interesting project.
A second reason for holding up our work on VLIW is that single-processor performance is not the bottleneck in todays systems. It was obvious to us that improving the memory subsystems of our computers was far more beneficial, and the first implementations of the new memory subsystems have already demonstrated this to be true. So far, we have talked only about single processors and how they might be incorporated into the AS/400e series. In the next section, I look at how multiprocessors will likely evolve over the next several years.
A major topic at every computer architecture conference I attend these days is scalable shared-memory multiprocessors. I certainly believe this form of multiprocessor will provide the future growth we need in our computer systems. There is much less interest in MPP designs, which do not share memory, because they are less general purpose and lend themselves to fewer types of applications. Besides, scalable shared-memory designs offer more challenges for us computer architects.
I introduced centralized shared-memory systems and distributed shared-memory systems in Chapter 2. A centralized shared-memory machine has a central memory with multiple processors sharing this memory. When people talk about SMP, this is the model they have in mind. Because in a centralized shared-memory machine the amount of time required for any processor to get to the central memory is the same, these are usually called uniform memory access, or UMA, machines.
A distributed shared-memory machine has the memory distributed among a number of nodes, where a node contains a small number of processors connected to the nodes memory in an SMP fashion. This definition of a node is the same one I used in Chapter 2, where a node has processors and memory but no disk or other I/O devices. The address space for all the nodes is shared, meaning that any processor can address memory in any node. A simple way to think about this is to picture a piece of the shared memory packaged with each node in the system and a high-speed global interconnection between the shared memory pieces. Each node has its own local memory bus connected to its piece of the shared main memory, but any processor can access any location in the shared memory across the global interconnection. The difference is the time required for the access. Local accesses will be faster than global accesses, which is why this cluster of SMP nodes is described as a nonuniform memory access, or NUMA, machine.
| Previous | Table of Contents | Next |