<html><head>
 <title>LinkSure 0.90</title>
</head>
<body bgcolor="#FFFFFF" text="#000000" link="#FF0000" alink="#FFCCA8" vlink="#FF5433" marginwidth="0" marginheight="0" leftmargin="0" topmargin="0">
<a name="topofpage"></a>
<table width="100%" border=0 cellspacing=0 cellspacing=0>
<tr bgcolor="#000000">
<td width="125" align="center" valign="center">
<img src="ls.gif" alt="LinkSure" width="29" height="27" border="0">
</td><td width="100%" align="center">
<font color="#ffffff" face="arial, helvetica" size="5">
<b>LinkSure</b> v0.90</font><br><font size="4" color="ffffff">HTML Link Checker for RiscOS</font>
</td></tr>

<tr valign=top><td bgcolor="#FFCC99" width="125">
<center><font size="4">Help for<br>
<b>LinkSure</b></font></center>
<hr>
<font size="2">
<img src="pad.gif" alt="-" width="125" height="2" border="0">
<br>&nbsp;<a href="#Disclaimer">Disclaimer</a>
<br>&nbsp;<a href="#Use">Use</a>
<br>&nbsp;<a href="#Options">Options</a>
<br>&nbsp;<a href="#Best">Getting the best...</a>
<br>&nbsp;<a href="#Problems">Problems</a>
<br>&nbsp;<a href="#Signing off">Signing Off</a>
<br><hr>

<br></font><center><font size="1">
Designed using<br>
<a href="http://www.goodwin.uk.com/richard/programs/html3/"><b>HTML&#179;</b></a><br>
</font></center>
<hr>
</td><td>
<ul>
<table border="0" cellpadding="2" cellspacing="2">
<tr><td colspan="2" align="center" bgcolor="#FFCC99"><B><I><font size="4">About this program...</font></B></I></td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Purpose:</B>&nbsp;</td><td>Check links in HTML documents</td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Author:</B>&nbsp;</td><td><a href="mailto:richard@goodwin.uk.com">Richard Goodwin</a></td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Status:</B>&nbsp;</td><td>Freeware
<P>This program is &copy; Rich Goodwin 1999-2001; it may be freely distributed if:
<ul>
<li>All files remain intact and unaltered
<li>No unreasonable charge is made
</ul>
This program uses the EasySockets module by Justin Fletcher.

</td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Version:</B>&nbsp;</td><td>0.90 (March 2001)<br><a href="History">History file</a></td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Features:</B>&nbsp;</td><td>Check local links, and sites on the Internet.<br>
Now with added non-http checks (e.g. newsgroups, telnet, email).</td></tr>
<tr valign="top"><td align="right" bgcolor="#FFDDAA"><B>Web:</B>&nbsp;</td><td><a href="http://www.goodwin.uk.com/richard/programs/">http://www.goodwin.uk.com/richard/programs/</a></td></tr>
</table><P>

<hr width="50%" align="center">
<h2><a name="Disclaimer">Disclaimer</a></h2>

<ul>
<li>If it doesn't work, tough.
<li>If it trashes anything, well, that's a good lesson in taking backups.
</ul>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>
<hr width="50%" align="center">


<h2><a name="Use">Use</a></h2>

Drag a HTML, WML, PHP and SHTML documents on to the <i>LinkSure</i> icon or main window and sit back!
<P>
<center><img src="window.gif" alt="Window" width="356" height="228" border="0"></center><br>
The <tt>Status</tt> area shows you what the program is currently trying to do; the <tt>Previous</tt> area shows the result of this for the previously checked link.  Under this is a counter showing elapsed time (although this only goes up after a link is checked, not during the actual checking).
<P>
There are also two buttons, <tt>Again</tt> and <tt>Cancel</tt>; <tt>Again</tt> will redo the check, which is useful if you've checked a page and made corrections (you don't have to drag the same file back in), and <tt>Cancel</tt> will either cancel the currently active page check (after it finishes with whatever link its checking) or it will close the main window if there is no checking going on.
<P>
<center><img src="eg.gif" alt="Example" width="329" height="359" border="0">
<P>
<I>An example results page</I></center>
<P>
The main bulk of the results page shows the URL being checked (with a little icon to show what type of URL it is), the line that this URL appears on in the original document, the HTTP result code of trying to fetch this URL, and an English translation of the result.  At the top of the page is a summary of errors - 404 (not found) and 500 (server error) problems are counted as major problems, and redirections (the fact that you could probably use a more direct URL) are counted as minor problems which you could probably ignore.  The <tt>meaning</tt> text is usually coloured so you can pick out errors quickly - ranging from green (OK) through yellow and orange (might want to check) to red (not found or server error).
<p>
<h3>Alternative link types</h3>
Previously the program ignored anything that wasn't in the form
<tt>href="http://&lt;some_stuff&gt;/file.ext"</tt>, 
<tt>href="../file.ext"</tt> or
<tt>href="file.ext"</tt>.  However, I've tried to add to the functionality so that a wider range of link types are covered.
<p>
New link types that are supported are anchors (<tt>href="#topofpage"</tt>), newsgroups, email addresses and telnet.  Anchors work by pre-processing your web page and doing a simple check for any <tt>name="..."</tt> or <tt>id="..."</tt> anchors being set up (up to 200); then when the program encounters a link to one of these anchors it checks its internal database to see if it is valid.  If the anchor is present all's well and good; if there's an anchor of the same name but with a different case, that'll be flagged as a minor problem; and if the anchor isn't found it'll come up as an error.
<p>
<h4>Newsgroups</h4>
Newsgroup links are tested by connecting to a CGI script on my web server and checking a news server to see if the group is valid; as such you need to be online to do this test, and have the relevant option switched on.  This was done in this particular way because an offline version would require a <i>huge</i> database of newsgroups (at least two megabytes), and an online check is much easier to achieve via a CGI than a desktop BASIC program.  It does mean that, for instance, if you're connected via a service other than the one my server's using you might not get a true result.
<br><b>Format:</b> I used to use <tt>news:group</tt>, but apparently the format is <tt>news://group</tt>.
<p>
<h4>Mail</h4>
<b>email</b> is checked for @ characters (too few or too many), and domains with no full stops in are flagged; bad links with @ in which could be email links without the mailto: on the front are also mentioned.
<p>
If you're using the online options, and the address didn't fail the above checks, then another CGI on my server will attempt to telnet into the mail server for that domain and figure out if the address is valid.  A positive or negative response is easy enough to figure out, but some servers don't give any usable response so I've made these return 302 ("ambiguous response" in this case, rather than a proper HTTP "redirected" message) - in these cases you're going to have to find another way to check I'm afraid.
<p>
For the techie I'm using RCPT TO rather than VRFY or EXPN, as the latter are usually disabled for security reasons; I also don't check fallback servers etc., but I do now look up the correct MX record for the mail domain rather than using the scattergun approach of v0.80.
<P>
<h4>Telnet</h4>
Telnet URLs are checked purely by trying to connect to the given server on the given port, or port 23 (telnet) by default.
<br><b>Format:</b> <tt>telnet://server:port</tt> (any filenames afterwards are ignored).  I thought <tt>telnet:server</tt> was fine, but it confuses some browsers, especially if you want to specify a port.
<br><b>Ports:</b>  Here's a few useful ports; if you're really weird you can use <i>LinkSure</i> to try connecting to servers on certain ports to see if the mail, ftp etc. daemons are running.
<ul>
  <br>21 ftp 
  <br>22 ssh (secure telnet)
  <br>23 regular telnet
  <br>25 smtp (mail sending)
  <br>79 finger
  <br>80 web server
  <br>110 pop3 (mail fetching)
  <br>119 nntp (newsgroups)
</ul>
This obviously opens up the idea that you might be able to use <i>LinkSure</i> as a crude port scanner - a quick FOR...NEXT loop to generate a page with telnet links to all the ports you want to check on a certain server, run it through <i>LinkSure</i>, and you get a list of openports back.  However, be aware that this sort of usage is not only morally dubious (unless you're the Sysadmin of the box of course), it may also backfire as many servers these days can detect port scans and even connections to ports where services aren't running and block all access from you.
<p>
<small>A cautionary note: while writing the the mail checker I forgot to specify that <i>LinkSure</i> should connect to the online part of the program on port 80 (web), so it defaulted to port 23 (telnet).  I don't have a telnet server running, so my server machine saw this as an attempt to probe it for nefarious purposes.  Suddenly my RiscPC locked up - the program that uploads what MP3 I'm playing locked up, my mail checker stopped, and I couldn't even fetch web pages such as <a href="http://www.iconbar.com/">The Icon Bar</a>.  I thought my network connection had stiffed, but I soon figured out what had happened - it was because I'd been locked out of my own machine!  Luckily I had a third machine to fix the problem, but it just goes to show what can happen when you mess around...
<p></small>

<h4>Other link types</h4>

<b>FTP</b> links are not currently checked, but the format is looked at - for example only <tt>ftp://</tt> is valid - <tt>ftp:/</tt> and <tt>ftp:///</tt> are flagged as errors.
<p>
I've also added support for root directories, so links in the form <tt>href="/index.html"</tt> can now be checked.  See the sections below for more information.
<p>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>

<hr width="50%" align="center">
<h2><a name="Options">Options</a></h2>
New in version 0.80 are some simple options, which can be accessed from the menu or by clicking with the right-hand mouse button (Adjust) on the <i>LinkSure</i> iconbar icon.

<p><center><img src="opt.gif" alt="Options window" width="281" height="160" border="0" hspace="10">
<p><small><i>Options window</i></small>
</center><br>
<ul>
<li><b>Check offline links only</b> will switch off any remote checks, handy if you're not connected to the Internet and just want to check links on your hard drive.
<li><b>Skip news check</b>.  Newgroup checks require <i>LinkSure</i> to connect to some CGI scripts on my website; you can skip this individually, or obviously if you're only checking offline links any checks that require an online checker will also be skipped.
<li><b>Skip mail server check</b>. <i>LinkSure</i> can try to connect to the SMTP port on the mail server at the domain specified in the email address; if this fails it'll try adding <tt>smtp.</tt> to the front, and if this fails tries <tt>mail.</tt> instead.  Obviously this requires Internet access, one or more attempts at a telnet connection and even then you can't be sure that the email address is valid, only that the domain exists and probably accepts mail.
<li><b>Root directory</b> sets where the program starts looking for files reference in the following way: <tt>href="/index.html"</tt> (i.e. starting with a slash).  The program can automatically detect some directory names as being probable roots, but it is safer and more accurate to set it yourself.
</ul>
Clicking <b>OK</b> will set the options, and <b>Cancel</b> exits the window without changing what the options were before you opened the window.  You've got to click OK for the options to be activated.
<p>
The options are automatically saved when you quit the program, and loaded when you start the program, so <i>LinkSure</i> should always start up with the same options as when you last used it.
<p>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>

<hr width="50%" align="center">
<h2><a name="Best">Getting the best out of it</a></h2>

<h3>Absolute links</h3>
One of the things I've added is checking for links in the form of <tt>href="/index.html"</tt>.  I didn't use a lot of these types of links when I first wrote the program because you can't follow them locally - if you load the file off your hard drive and have a link or image referenced in this form, the web browser won't know where to find it (which is exactly the same problem <i>LinkSure</i> had).  However, for big websites with a standard navigation layout this type of link is a must, unless you want to change every page by hand (or use server-side includes for navigation).
<p>
To get around the problem of not knowing where "/" refers to, <i>LinkSure</i> needs to know the directory on your hard drive that "/" equates to.  This can be achieved in two ways - either set the root in the <a href="#Options">options window</a>, or let the program check the path of the file you drag in for you.
<p>
If you're using the automatic version, <i>LinkSure</i> will look through filename of your webpage for instances of <tt>site.</tt> and <tt>pages.</tt> - directories ending in <i>pages</i> or <i>site</i>.  The check is case insensitive, and you can have other characters on the start of the directory name, such as <tt>WebPages</tt>.  The program then works out how to get from your file to this directory, and once it has got there it checks the link as normal.  If you want to use this method however, you might need to move your website to a new directory, which might mess up programs that check to see if any files have been updated.
<p>
All of this is a little difficult to explain clearly in words, so here's an example from my own hard drive:
<br>&nbsp;
<center><img src="dirs.gif" alt="Directory listing" width="421" height="299" border="0" hspace="10">
<p><small><i>Directory listing</i></small>
</center><br>

I have quite a few sites on my drive (only a few of which are shown), so I have a <b>WebSites</b> directory; inside that are named directories for each site, for instance <b>Mabel</b> for <a href="http://www.houseofmabel.com">www.houseofmabel.com</a>.  Inside that I have some other directories - usually stuff like <tt>Raw</tt> and <tt>Dump</tt> to hold the ArtWorks and Photodesk files I used to create the site just in case I need to make any changes - and the important one for <i>LinkSure</i>, one called <b>webpages</b> where the actual site mirror is kept.  It is this directory that <i>LinkSure</i> will assume is the root of the server for a <tt>href="/index.html"</tt> type check.
<p>
Although the <b>WebSites</b> directory won't match for <i>LinkSure</i>'s purposes (it has an extra "s" on the end), be aware that <i>LinkSure</i> will check until the <u>last</u> instance of <tt>site</tt> or <tt>pages</tt>, so if you have <tt>ADFS::4.$.Website.Data.Website</tt>, it is the second instance of <tt>Website</tt> that the program is interested in.  This does mean that if your site legitimately has a <tt>pages</tt> directory in it (e.g. <tt>http://www.domain.com/book/pages/example.html</tt>) <i>LinkSure</i> will incorrectly assume <tt>example.html</tt> is in the root of the domain if it is checking these types of links.  To get around this you'll have to set the root directory manually in the options window.

<p>
<h3>Verifying links "by hand"</h3>
As with any machine checks, to be absolutely sure of a problem you should really try to verify things by hand, especially the major errors.  The program is getting better and better at giving good results, but it is best to be safe than sorry.  This is why some things that might not be considered errors but can't be properly verified - for instance email addresses - still get flagged as potential problems.
<p>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>

<hr width="50%" align="center">
<h2><a name="Problems">Problems</a></h2>
<ul>
<li><B>It just sits there.</B><br>
Some sites just don't respond to requests for information, but the new version should time out after about 30 seconds to a minute.
</ul>
<P>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>

<hr width="50%" align="center">
<h2><a name="Signing off">Signing off</a></h2>
<B>Rich Goodwin</B><br>
<a href="mailto:richard@goodwin.uk.com">richard@goodwin.uk.com</a><br>
First published Thursday 4th November 1999
<P>
<a href="#topofpage"><img src="up.gif" border="0" alt="^" width="16" height="16"></a>&nbsp;Back to top<P>

</ul>
</td></tr></table></body></html>