| Previous | Table of Contents | Next |
In Chapter 2, we introduced the SMP support on the AS/400, where all the main processors are operating concurrently against a shared memory. For most types of processing, each processor is working on a different job in the system, sharing the memory as required. For database processing, each main processor could be working on a part of the same job. The parallel-database feature of DB2/400 does exactly that. A single query is broken down into separate, independent queries, and then these separate parts of the original query are run in parallel across the multiple main processors in the system. This arrangement can significantly increase the performance of the original query. The CHGQRYA (Change Query Attributes) command includes an option to let the user specify that the query is to take advantage of the multiple processors.
The types of query functions that will benefit from this enhancement are table scan, group-by, index scan, and join. Some functions internal to SLIC, such as index build (discussed in the Machine Index section of this chapter), also benefit from the SMP parallel database support. These enhancements are available on all V3 and V4 systems, except for the parallel index build, which is a function in RISC systems only.
The AS/400 also supports MPP configurations. To create an MPP configuration, multiple AS/400 systems are connected together with high-speed links to form a cluster of systems. One way to link the systems is through the use of a fiber-optic connection product called OptiConnect. On the AS/400e series, the SAN interconnection can be used for even higher data rates between systems. With the database spread across the disks on each system in the cluster, extremely large databases can be built with several hundred processors working in parallel.
IBM calls this MPP configuration a loosely coupled parallel database system, because outside of an individual system in the cluster there is no memory sharing. This is the shared-nothing approach I described in Chapter 2, and it is similar to the approach used in the SP2. The difference is that the AS/400 nodes in the cluster are packaged in separate physical boxes. No matter what packaging is used, to the user the cluster appears as a single database.
Loosely coupled parallel database support lets queries be broken into smaller units of work that each node can resolve. The difference from the SMP parallel database support is that here each node has its own memory and disk space. Each node in the cluster has a portion of the physical file or table. The entire query is processed on each node against the portion of the file resident on that node. Because each node is simply an AS/400, a node may contain one or more processors.
An application on any one of the systems in the cluster can access the entire database by simply opening it as if it existed entirely on the local system. DB2/400 makes this accessibility transparent to both the application and the end user. New CL commands have been added to name all the systems in the node group, and new parameters have been added to certain commands to allow distribution of files in the database across the nodes. Once distributed across the nodes, a file acts like a local file with respect to insert, update, and delete operations.
A major advantage of the loosely coupled parallel database support is that there is no upper limit to the number of nodes that can ultimately be interconnected, which means performance and capacity scalability in the future is essentially unlimited. In Chapter 11, we look at how the concept of AS/400 clusters might be extended in the future.
As we will examine in more detail in a following section, relational databases are organized as two-dimensional tables. An MDD has one or more additional dimensions. For example, suppose we want to analyze sales revenues by product, by geographic regions, and by time. We can create a three-dimensional data structure with a list of our products along one axis; time in days, weeks, or months along the second axis; and geographic regions along the third axis. When we are finished, we will have created a cube that looks very much like a three-dimensional spreadsheet with revenue amounts in each cell. We then can use various analysis tools to view sales by product and by region over a certain time period.
The AS/400 supports multidimensional data structures directly through database designs using DB2/400, or through products available from AS/400 business partners. The advantage of the multidimensional data structures is their capability to quickly answer business questions by slicing the data along any dimensions or drilling down through the structure to new levels of data. Because the response times to business questions are typically very fast, this multidimensional analysis is often called on-line analytical processing (OLAP).
Sometimes the various departments within an organization want to view the informational data in different forms. We can build data marts, which are smaller departmental warehouses, within the MDD. A data mart contains the informational data tailored to the needs of a specific department or work group. The data warehouse for the total organization then comprises a collection of these data marts.
The term business intelligence is used to describe the discipline of developing information that you then can use to drive business decisions. Business-intelligence tools are software packages used to analyze the data in an AS/400 data warehouse. These tools are typically PC-based and can directly access the AS/400 data warehouse. The three major categories of business-intelligence tools are
DSS tools let an end user develop a hypothesis and then create questions in the form of queries to test the validity of the hypothesis. A DSS tool assumes the end user has some general idea of what to look for in the data warehouse, and it enables the end user to build ad hoc queries and generate reports. This is the simplest type of tool because it just needs to return information that satisfies the question being asked.
EISs combine decision-support tools with some extended-analysis capabilities. Typically, access to facilities outside the data warehouse is also provided. For example, you can use some of the online news services available over the Internet to provide information about business markets around the world. As with the DSS tools, EISs assume the end user has some idea what to look for when creating a query or when drilling down into the data to prove a hypothesis.
| Previous | Table of Contents | Next |