

Word Hound Extras
=================

This section details the WHExtras, which allow the user to create new, and
also edit the existing, data files for use with Word Hound. WHExtras is a
set of 4 tools which allow you to :

1) Uncompress compressed data and index files into plain text files.

2) Automatically create 'thesaurus' index files.

3) Compress plain text data and index files for use with Word Hound.

4) Format the 'Jargon File' for use with (3)


These three tools are non-multitasking tools, and expect to find the files
they are to work on in the same directory as they reside. It is NOT advised
to run these tools in task-windows.

Before detailing how to use each of these tools, it is important that you
understand what files are what.	

 Compressed data files are files that contain dictionary and thesaurus
  entries. The must end with '-d' (eg. 'dict-d' and 'roget-d').

 Compressed index files contain indexes into the compressed data files.
  These must end in '-i' (eg. 'dict-i' and 'roget-i').

 Quick Index files contain indexes into the compressed index files. 
  These must end in '-q' (eg. 'dict-q' and 'roget-q').

All three of these must be present, and correctly formatted. For this reason
the three tools are supplied so that the user does not need to understand
these files, as all three will be created automatically. In this document
the term 'data suite' refers to a collection of these three files.


Uncompressed file formats
-------------------------

So that compression, and automatic index creation, can take place the
uncompressed data and index files need to be structured. This structure is
simple to understand and unrestrictive.

- Data file

First of all, the data file. The data file is the file which contains the
dictionary/thesaurus/what-ever entries. Each entry must be given a number,
and no two entries within one data file can have the same number. On the
line before each entry in the data file, simply place (on a line of it's
own) a hash ('#') followed by that entries number, eg.

#1

Each entry can contain as many lines as you like, but each line must be less
than 80 characters long. For consistency it is advised not to place blank
lines before, or after each entry. An example data file is supplied with the
tools, and is called 'EXdata' and this file also contains extra information
about entry numbers, etc.

For normal dictionary type entries that is all you need to know. However for
thesaurus entries there are some extra rules. Each entry in the thesaurus
needs to start with a number, and this number does NOT need to be the same
as the entry number. This number then is followed by a '.' and this in turn
followed by a 'key entry' which in turn is followed by a '.'. This 'key
entry' is the word (or phrase) which is displayed in the 'Selection' window
of Word Hound, and gives a general description of the whole entry.

Next each word or phrase within the thesaurus entry needs to be separated by
one of the following punctuation marks : ',' (comma), ';' (semi-colon) or
'.' (full stop). (Note the word (or phrase) within an entry also needs to be
followed by one of the punctuation marks.

Here 570 is the data file entry number, 566 is the thesaurus entry number,
and 'Phrase' is the key entry. Note that the phrase 'turn of expression' is
split across lines, this does not matter.


- Index file

The index file is simply a list of indexes. There is one index per line, and
each index is made up of a word followed by a tab character followed by the
data file entry number, for example :

phrase[09]570

where [09] represents the tab character. The index file does NOT need to be
sorted into any particular order as !Bulder will automatically sort the
entries into the required order, however it is advised to make sure that all
indexes are lower case. An example index file is supplied with the tools,
called 'EXindex'.

You are allowed a wide variety of characters within indexes, which are
treated in the same way as 'words' within Word Hound. A 'word' is defined as
a continuous sequence of characters excluding the following :

   space, comma, full stop, semi-colon, colon, exclamation mark, ampersand
   (&),  	double quotes ("), any brackets (parentheses, curly, square or
   angled) and any 	character below space within the Ascii table.

Word Hound will allow any character above 127 within words (and thus
indexes), but support is limited in that there is no concept of 'case' for
characters above 127, thus care should be taken for foreign (non standard)
characters within any data files, and in particular it is advised to only
use lower case versions.


The tools in detail
-------------------

1) !Extractor

!Extractor takes a compressed data file (-d file) and a compressed index
file (-i file) and produces plain text versions of each. The quick index
file (-q file) is not needed by !Extractor.

When you run !Extractor it will ask you for the file name of the 'data
suite' you wish to extract. This is the name of the compressed data file
less the '-d' ending (or the name of the index file less the '-i'). Both the
'-d' file and the '-i' file must be in the WHExtras directory.

!Extractor will then ask for the name of the file to save the uncompressed
data to, and then the name of the file to save the uncompressed index data
to. When you have entered both these file names, !Extractor will perform
extraction.

The data and index files produced will be of the formats defined above. The
data file entry numbers will start from 1, and the produced index file will
be alphabetically ordered.


2) !MakeIndex

!MakeIndex takes an uncompressed data file (of the format described above)
and produces a thesaurus index file for it. To be able to work the data file
needs to comply with the rules defined above.

!MakeIndex first asks you for the name of the data file to use, and then the
name of the file it should use to output the index to. Once you have entered
these, !MakeIndex will create the index file.

If any of the data entries have a mis-matched number of brackets you will be
warned, giving the data entry number for the offending entries. (Note, words
contained within brackets are not included within the index file).

!MakeIndex will create index entries for all single words in the data file
(not phrases) apart from some special cases, namely N, V, Adv, Adj and Phr
which are used in the thesaurus to denote parts of speech. All words will be
converted to lower case automatically, and each word will only be added once
for each data entry number.


3) !Builder

!Builder takes an uncompressed index file, and an uncompressed data file and
produces the three files needed to be able to use them with Word Hound.

When you run !Builder you will be asked to enter the file names of the data
and index files (the uncompressed ones). It will then perform all the
necessary functions to produce the three files, which it will call :

     new-d     new-i    and    new-q

!Builder keeps you informed of what it is doing at every stage. These stages
are as follows :

Reading the index file - It needs to do this twice, once to simply count the
entries (it tells you this number) and then to actually read in the entries.

Scanning the data file - This is needed to build up the compression
algorithm to be used to compress the data. Each data file has it's own
compression algorithm to give it an optimal compression.

  Compressing data     - This creates the 'new-d' file.

  Sorting Indexes      - All the indexes need to be sorted into alphabetical
                         order to create the index files.

  Creating index files - This creates the 'new-i' and 'new-q' files.

All you need to do now is to decide of the name to give the data suite,
change the 'new' part of the file names to match that, and move the files to
the directory where you keep your dictionary/thesaurus/what-ever files. If
you have created a completely new data suite you will need to inform Word
Hound about it (see below).

4) !JMake

The 'Jargon Lexicon' is a file (available from various sources) which forms
the text for the book "The New Hacker's Dictionary" (edited by Eric Raymond,
published by The MIT Press). This book is a comprehensive compilation of the
remarkable slang used by today's computer hackers.

The file version of the book is a plain text file, and is freely distributed
around the academic (and other) networks. As such it forms another source of
data for Word Hound, this is where !JMake comes in. It takes the lexicon
body of the data files and produces the uncompressed data file and
uncompressed index file needed by !Builder.

To use !JMake you first need to edit the 'Jargon lexicon' file to strip away
the introduction and appendices, thus the first line should now start "= A
=" and the last line should be the last line of the final definition. This
edited version should now be saved (as "Jargon") inside the WHExtras
directory. If you now run !JMake it will then create 2 files, "jdict" and
"jindex", which can be passed to !Builder in the normal way.


The 'Data Suite' list in the Word Hound Startup file.
-----------------------------------------------------

So that Word Hound can keep up to date about what data is available to it,
version 1.16 uses a special file (called the 'File-list' file) which details
each data suite available. This is a simple text file, which on it's first
line should contain the total number of data suites available, and then on
subsequent lines (one line each) each data suite is listed.

For example, here is the relevant section of the default 'Staup' file (which
is called 'Startup' and resides within the application directory) :


This says that there are 2 suites available. The first is called 'dict' (ie.
it should look for 'dict-d', etc.), it is a dictionary file (denoted by the
'1'), and it's full name is 'Dictionary'. If you place a '!' before any of
the listed entries, then the data suite defined by that line will start up
'deselected' within Word Hound.

The number given as the second argument is important as it tells Word Hound
if the data files should be scanned for quick definitions, and it also tells
it a bit about how the data is formatted, so it can find something useful to
display in the Selection window (third column).

Data Types
----------

0 - Thesaurus data. This is formatted as described above. Word Hound takes
    all text from the first non-space character on the first line upto the
    second '.' as the text to display in the selection window.

1 - Dictionary data. This is data as formatted in the main dictionary. This
    data has the format (on first line) of an optional number, followed by
    the word that the definition refers to, followed by a pronunciation
    guide (enclosed in '\'s). Word Hound will take the text between the
    optional number and the start of the pronunciation guide as the text for
    the selection window. Alternatively it will take the text upto the first
    ',' on the first line of the definition.

2 - Jargon file data. This refers to data not supplied with this application
    but available from various sources. The text used by Word Hound is
    contained between ':'s on the first line.

If any other data type number is given, Word Hound will simply repeat the
word in the second column in the third column of the selection window.

All non-zero data types are treated as definition types, and as such may be
used by the quick-definition.


Memory requirements
-------------------

Each tool has it's own memory requirements, but the one which requires the
greatest amount of memory is !Builder. This is because it needs to keep the
whole uncompressed index file in memory whilst it works. This can mean that
you run out of memory when trying to build a data suite. This, however, is
not a major problem.

If you do run out of memory the simple solution is to split a data suite
into several parts and build each separately. So, for example, if you wanted
to add to the dictionary, but can not edit the (over) 3 meg. uncompressed
data, simply decompress the data, then split it into several smaller files
(making sure you don't split in the middle of an entry), and also split the
index files accordingly. Then you can simply build each separately, giving
each a different data suite name (eg. dictAJ, dictKQ and dictRZ to depict
their ranges, but this is just an example). Then if in your 'File-list' file
you give each of the new data suites the same full name (this is allowed)
the user will be none the wiser.

Actually there is no real reason to ever need to edit the supplied
dictionary data suite (unless you wish to make corrections), as you could
simply create a supplementary dictionary.


IMPORTANT NOTE
--------------

Because of the internal workings of the compressed index files, no
compressed data file can be more than 2 Meg. in size. If the compressed data
file is bigger than this limit, the index file will not be correct for
entries that lie beyond that limit.

If a data file starts to get near that limit, it is a good idea to split it
into smaller parts (as described above) or start a new data suite.

Note that the supplied dictionary file is already near that limit, and so it
is not advised to add entries to that data file. A supplementary dictionary
is advised as it requires less effort than splitting it into smaller parts.

