| Previous | Table of Contents | Next |
As the new effective addresses are translated to virtual addresses using the segment table in memory for the new process, the contents of the SLB registers are updated one at a time. Using the memory table until the SLB registers are reloaded can add many processor cycles to every memory access. As a result, the performance of the process is degraded until some number of SLB registers have been loaded. If, for example, the new process switched in was our sequential database read, we would have to translate at least four addresses one each for the program, index, cursor, and data space just to get started. In reality, because each of these objects has multiple segments (each with its own virtual address), we would probably have to translate 8 to 12 addresses using the memory tables before the SLB registers would be of much use to speed up the translation process.
We just saw that the AS/400s single-level store does not use the segment table. The effective address is the same as the virtual address no mapping is required. All the virtual memory can be directly addressed from the program. The PowerPC processor hardware bypasses both the segment table and the SLB registers for AS/400 programs. This means there is no performance degradation due to the segment table on a process switch. Bypassing the segment table and the SLB registers can be a significant performance improvement for the AS/400, but it is only the beginning.
Later in this chapter, we will see in detail how a virtual address is translated into a real address using a page table. Depending upon the specific operating-system design, each user process in a conventional system may have its own private virtual memory with its own unique page table. An example of an operating system like this is Microsofts Windows NT. The memory management component in Windows NT provides a large, private, virtual address space for each process. This means a process switch not only has to change the page table to correctly map the new processs private address space, but it also has to purge all the TLB registers. After the process switch, addresses are translated from virtual to real using the new page table, and the contents of the TLB registers are reloaded one at a time. Just like reloading the SLB registers, reloading the TLB registers after a process switch degrades performance.
Only one virtual memory exists in the AS/400, so only one page table exists that everyone uses. Consequently, there is no need to purge the TLB in an AS/400 on a process switch. Because the TLB registers are not purged, an AS/400 also can make more efficient use of a larger TLB than some other system can. The TLB registers hold the most recently used entries in the page table. As time goes on, older entries are displaced by more recently used entries. With more registers, the likelihood that a virtual address that was translated in the distant past is still in the TLB is higher. When a process that ran in the past is again switched in, the larger TLB means some or all of its addresses may still be available. This would not be the case if the TLB registers had to be purged on each process switch. Once again, the single virtual memory of the AS/400 saves a great deal of processing time.
Finally, modern processors dont fetch and store information directly to or from the memory. They use cache memories. A cache memory contains portions of the main memory and has its own directory. Depending again on the design of the cache memory a particular computer uses and whether bits from the virtual address are used in the cache directory, the cache may have to be purged on a process switch. Again, however, this is not a problem for a single-level store.
Operating systems designed to work with conventional virtual memories usually try to avoid doing too many process switches, because doing so requires such a large overhead. When these systems are faced with a situation that requires lots of process switches, they must rely on high-performance processors to achieve acceptable system performance.
Process switching in an AS/400 is extremely fast compared to other systems, because there is so much less to do. The performance degradation when a new process starts is also less in an AS/400. As a result, the parts of the operating system, both OS/400 and SLIC, are designed to do lots of process switches. A few years back, the IBM Research Division did a study of the AS/400. This group found that, in a typical AS/400 user environment, a process switch took place about every 1,200 instructions. This was unbelievable to them, because some operating systems can take 1,000 instructions or more just to perform the process switch. Not so for an AS/400.
Because of this capability to rapidly switch between processes, the AS/400 excels in interactive performance. The System/38 and the original AS/400 were optimized for interactive, transaction-based applications. This optimization means many terminals can be attached to an AS/400. A single large AS/400 easily can support many thousands of concurrent users, something its competitors, such as Unix and Windows NT, have difficulty doing. In an application environment that requires many process switches, an AS/400 can outperform systems with faster processors, because it executes fewer instructions.
If we look at an application environment that does not require many process switches, we see a big performance difference. Consider, for example, a batch environment, in which a single process keeps executing for a long period of time. Here, the speed of a process switch does not play a significant role. Early AS/400 systems were not strong batch performers, because the processor performance showed through. In the past, the AS/400 never had high-performance processors.3
3Before some of my performance friends, such as Rick Turner, point it out, batch performance depends on more than just processor performance. Such things as memory sizes and disk I/O capabilities play major roles in batch performance.
Even in quoting performance for various models of the AS/400, in the past Rochester never wanted to quote processor ratings such as millions of instructions per second (MIPS) or MHz (just as MIPS measures how many instructions a processor can execute in a second, MHz measures the number of cycles per second the processor can execute). If two different processors must execute the same number of instructions or cycles to get a particular job done, and everything else is equal, then MIPS or MHz might give some indication of how the two processors perform. But if the processor doesnt have to execute as many instructions to do the same job, then neither MIPS nor MHz has any value.
| Previous | Table of Contents | Next |