| Previous | Table of Contents | Next |
Both EISs and DSS tools assume you know what to ask of the data warehouse. These are validation-driven decision-support tools that enable a computer search for information that validates some hypothesis. What if you dont know what questions to ask? You know there is valuable information on trends, patterns, and correlations hidden in the data, but you are not sure how to get at it. What you need is a discovery-driven decision-support tool. What you need is data mining.
Data mining is the discovery of information with little or no direction from the user. With data mining, the system searches through the data to determine patterns and associations. For example, a retailer may use data mining to determine an affinity between products by analyzing customer buying data. This capability to closely analyze and profile customer buying habits can be very valuable when you are creating product promotions or targeting specific customers with the best buying potential. Industries such as banking also are using data mining to detect credit-card fraud. By sifting through large collections of data, they can identify deviations from the norm using the data-mining tools. The possibilities are endless.
IBMs data-mining tools combine neural networks with statistical algorithms to find patterns and relationships in business data. A neural network supports the core data-mining step in the knowledge-discovery process. Data mining is a technology that came from the world of artificial intelligence. The neural-network technology IBM uses was first developed in a Rochester advanced technology group formed during the Fort Knox era (see the Appendix for more information about this group). The neural-network technology first appeared as an AS/400 utility in the early 1990s, but it now forms the basis for all data mining in IBM.
Metadata is data about data. This metadata is used to manage the data warehouse. The two forms of metadata are technical data and business data. The technical data contains the description of both the operational database and the data warehouse. The two descriptions enable the movement of data from the operational database to the data warehouse.
The end user uses business data to find information in the data warehouse. Think of this business data as a catalog of all the information in the data warehouse, including specifics about how current the information is and where it came from. This business data is also presented to the end user in business terms, so that the complexities of the underlying physical database do not show through.
Now that we have looked at some of the ways businesses are using the new database technologies on the AS/400, we can examine the fundamental concepts of DB2/400. Lets first look at how this remarkable database came to be.
The first commercial database with relational capabilities appeared in the System/38. It predated other relational databases in the industry by about three years. This unique database helped set the System/38 apart as a very advanced system, and many wondered where it had its origins.
The System/38 databases original developers were looking for a more efficient way to process records than the System/3 used. The first System/3 was designed as a unit-record machine. It could do only batch processing, which meant that an application would process all the records in a file, one at a time. The first records were on punched cards, and a deck of punched cards made up a file. Later, files could be stored on disk, although they were still processed the same as the cards had been.
A typical unit-record application would first sort the records in the file. The records would have multiple fields containing such things as customer name, account number, part number, and the like. One of these fields, called the key, was selected, and all records would then be sorted into some sequence based on the value in the key field. The mechanical card sorter in most unit-record installations was usually kept quite busy. After the file was sorted, it would be processed, one record at a time, until all the records in the file had been exhausted.
Interactive processing was later added to the System/3. The use of disk technology meant individual records could be accessed in random order. An index was used to identify the record to be accessed. An index is a small file where each record in the file has only two fields. The first field contains a key value, and the second field contains the disk address of the record that has the matching key value. A sort program was used to sort the index entries according to the key values. The index was then stored on the disk along with the file itself.
To find a record that had a certain key value, the system would first search the index. When it found the key value, the disk address stored with that key value in the index was used to fetch the full record from the file. Because the System/3 had fairly small memory sizes, it was not possible to store much of the index in the memory. This made searching the index fairly inefficient because several disk accesses were required.
The first member of the System/3 family to be designed as an interactive system, rather than a batch system, was the System/34. The System/34 also had relatively small memory sizes, so IBM needed a better way to search for an index entry. The approach it took was to eliminate the requirement to fetch the index from the disk.
Certain tracks on the disk were reserved for the index and special hardware was designed into the disk controller. The desired key value was passed from the processor to the disk controller. The disk controller then began to read the information on the track, looking for the key value. When the controller hardware saw the key value, it read in the next address field and passed this back to the processor. The processor then used this address to fetch the actual record from some other part of the disk.
This type of disk operation was called a scan. The scan function greatly improved the efficiency of interactive processing by completely eliminating a disk access to fetch the index. However, a smaller index that could fit into memory did have to be built over the index on the disk. The index in memory identified which track on the disk to search. This same scan operation was later implemented in the System/36 to do file processing.
The System/38, too, needed to do interactive processing very efficiently. The disadvantage of the scan implementation is that it tied file searching very closely to the hardware. There were also other limitations in terms of the number of indexes possible and the ways they could be processed. Because the System/38 was to have a single-level store, the developers decided to put all files and indexes into this large memory.
| Previous | Table of Contents | Next |