Previous Table of Contents Next


A RISC processor looks at the next three to six instructions in the instruction stream and dispatches them to as many functional units in the processor as are available. The RISC compilers are responsible for scheduling independent instructions in that instruction stream, so as many as possible can be dispatched by the processor each cycle. The back end of the compiler takes the intermediate text and generates the binary-level machine instructions. To schedule these instructions for the processor usually requires examining the generated instructions from a range of intermediate text to find enough independent ones to schedule.

In a VLIW machine, the number of functional units increases dramatically to achieve more parallelism. A number in the range of 16 to 32 functional units is not unreasonable in the near term, with even larger numbers possible later on. To schedule instructions for this kind of machine requires the back end of the compiler to examine a much larger range of intermediate instructions. The compiler may have to look at the generated code from hundreds or even thousands of intermediate instructions to find 16 to 32 independent binary-level instructions to be dispatched each cycle. The dynamic, on-the-fly approach looks at only a few instructions at a time as it generates the VLIW instructions. The big question in the industry was how efficiently this on-the-fly translation would be able to schedule work for the many functional units.

Recent information from Intel seems to suggest that Merced will be more akin to the superscalar RISC-like approach used in the Pentium Pro. This design may still use some of the basic concepts of VLIW, including parallel dispatching of multiple instructions. One Intel executive said they took the VLIW work at HP and the RISC/CISC work at Intel and came up with something new. He said this new kind of architecture is beyond RISC and beyond VLIW. In Rochester, we are watching with great interest.

The History of VLIW in Rochester

In Rochester, the interest in long-instruction-word machines started in the early 1980s. At that time, we established an advanced-technology organization in the Rochester laboratory to investigate new technologies that we might someday use in our future systems. My role in that organization was to manage the systems group.

In the systems group, we concentrated on high-performance computing and parallel processing, mostly with the System/38. We were convinced that some day we would build very high-end models of a System/38, and we wanted to be ready for that day. One area Roy Hoffman led was the exploration of coprocessors. The idea was to add special-purpose processors to the System/38 to efficiently execute applications for which the System/38 was not well suited. One of the coprocessors we attached to a System/38 was a high-performance, floating-point engine. The goal in an advanced-technology area is not necessarily to do the actual design to be shipped to customers, but to demonstrate that the ideas are sound and the design can be built. And the best way to demonstrate an idea is to build a working prototype. As long as we were at it, for the floating-point project, we decided to go all out and build a System/38 with the performance of a supercomputer. Our internal goal was to demonstrate applications with heavy floating-point usage running on our prototype that would outperform a System/390 with its vector feature. We set about achieving our goal with the help of a coprocessor from Floating Point Systems (FPS).

FPS in 1975 introduced the AP-120B, which was the first member of FPS’s array-processor family and was used primarily in signal processing. Array processors operate on ordered sets of data, usually vectors or matrices. In 1980, FPS introduced the FPS-164 to extend the AP architecture into large-scale scientific processing. The FPS-164 was a full 64-bit processor designed for this type of scientific computing. It could hold its own against any of the supercomputers of the day, including the Crays.

The FPS processor was not a standalone system; it would always be attached to some other host computer system. Because we thought it would make a great coprocessor for a System/38, we bought an FPS-164 and attached it to a System/38. We also began to track down commercial applications that needed high-performance, floating-point computations. We wanted to show that this type of computing was applicable to more than just scientific work. We found several applications, the most promising in the banking and securities industries.

The FPS-164 was not exactly a small system. Physically, this coprocessor was much bigger than the System/38. It also had a few environmental requirements that the System/38 didn’t, such as lots of cold air. We built a special room with a raised floor and some of the biggest air conditioners our maintenance people had ever installed in such a room. The fans in the FPS box drew cool air from under the floor. When we turned the unit on, it sounded as if a very noisy hovercraft was in the room with us. When it was off, the room was too cold for any of us to work. For a computing engine, however, the FPS-164 was really fast.

FPS had plans to bring out newer technology versions of the FPS-164 that would operate in an office environment and easily fit alongside the System/38. But because the Fort Knox project was cancelled and we needed to put all of our energies into Silverlake, we never had a chance to add one of these to the System/38. We did learn several things from this work, including the use of coprocessors for AS/400s.

The FPS systems were among the first long-instruction-word machines that contained multiple operations per instruction. The machine had 10 functional units, and each unit needed its own subinstruction every cycle. The long instruction word had separate subinstructions for each of the 10 functional units. A complete vector loop was possible in a single instruction.

Instead of using an optimizing compiler to create the long instruction words, the machine provided assembly-language libraries to make it efficient to use. The host computer handled the program logic and called out the long-instruction-word routines to be run on the FPS machine. There were similarities between this type of programming and wide-word microcode such as the HMC used in the System/38. Although our HMC instructions did not have as many bits as the long instruction word in the FPS, each HMC instruction did start multiple functions in the System/38 processor. For a time, we looked at how we could incorporate into our HMC some of the techniques used to dispatch the multiple functional units.


Previous Table of Contents Next

Copyright © NEWS/400 Books