Previous Page TOC Next Page



- 21 -
Composing and Editing HTML Pages
by Dick Oliver and Mark Thomas


This chapter gives you a fast-paced but friendly introduction to creating your own Web pages. After you learn the basics here, turn to Chapters 23 through 26 for comprehensive coverage of all the fancy tricks you can use to make pages especially for Internet Explorer 3.0. These chapters also show you how to ensure that your Explorer-enhanced pages are completely compatible with all other major Web browser software, including Netscape Navigator.

What Are HTML and HTTP?


Web pages look fancy when viewed with a browser, but they're just plain old text underneath. To tell the browser where to put headings, graphics, and other visual emphasis, you mark up the page with special codes called HTML tags.

HTML stands for Hypertext Markup Language. It is a computer language being developed by members of the World Wide Web Consortium at the Laboratory for Computer Science at the Massachusetts Institute of Technology with input from Internet users and companies around the world. HTML is one particular application of a much broader text-encoding language called Standard Generalized Markup Language (SGML). SGML's purpose is to unify the exchange of electronic text-based information across all computer network operating systems. The specific purpose of HTML is to facilitate the interpretation and presentation of not only text content but also multimedia-based information among computer networks.

The universal machine language that enables all these computers to communicate among themselves over the World Wide Web is known as the Hypertext Transfer Protocol (HTTP). HTTP is the backbone of the World Wide Web. It is the communications protocol different computers across the Internet must use if they are to understand each other and be allowed to exchange information over the Web. For the most part, if you are interested in creating Web pages and making information available to the Internet community, your understanding of HTTP does not need to go deeper than simply knowing what it is and that it provides the engine for HTML functions.

HTML has become an integral component of the World Wide Web, a project developed at The European Laboratory for Particle Physics (also known as CERN). HTML documents are virtually always interpreted by software known as Web browsers. Microsoft Internet Explorer, Netscape Navigator, NCSA Mosaic, and Lynx are all examples of Web browsers. The vast majority of people accessing the Web today use either Microsoft Internet Explorer or Netscape Navigator, which both support very similar extensions to "standard" HTML.

HTML documents typically contain text and images; in that sense, they are conceptually similar to ordinary magazine articles or term papers. The difference, however, is that HTML goes beyond the limits of ordinary text by including hypertext references to other documents. Text from one document can, if the document's creator so chooses, be highlighted to indicate that there is more information available on this particular subject. Several documents can be hyperlinked among each other in this way, or a single document can contain references within itself.

Hypertext references, also known as hyperlinks or Uniform Resource Locators (URLs), combined with a variety of multimedia resources, text, and content-enhancing tags are what give the Hypertext Markup Language its unique capacity to communicate and enrich the presentation of document content and visual information.

Choosing an Editor


Any standard text editor or word processing program can create HTML documents. Basic text editors such as Windows Notepad or the Macintosh's SimpleText are almost always included with any computer you purchase. These "no-frills" word processing programs do not have all the features of the most popular word processing programs such as MS Word or WordPerfect.

You also can find similar programs with enhanced capabilities on the Internet. MegaEdit is an excellent text editor for Windows, and BBEdit is an excellent enhancement for the Macintosh's SimpleText editor.

When you are using commercial word processing software to create HTML documents, be careful how you save your documents. When you use Microsoft Word, for example, save your HTML documents as Text Only documents and not as Word for Windows documents. Saving a document as a Word for Windows document means that that document contains formatting and other elements that are unique to Word for Windows. Because the Hypertext Transfer Protocol assumes that HTML documents are generic ASCII text, the presence of those formatting elements that are unique to Word for Windows will produce unpredictable results when an HTML user agent attempts to access them. (Microsoft Internet Assistant enables you to save formatted HTML documents from within Word and other programs. You'll find out more about how to do this in Chapter 25, "Publishing with MS Internet Assistant.") Another very important fact to keep in mind is that HTML user agents do not always recognize a document as an HTML document unless its file extension is .htm or .html.

Although you can use any text editor or word processor to create HTML pages, you'll probably enjoy the convenience of a dedicated HTML editor program if you do very much Web publishing. These programs come in two varieties: those that show HTML tags as plain text and those that attempt to offer a "What You See Is What You Get" (WYSIWYG) preview of pages as they will look when viewed with a Web browser.

Figure 21.1 shows the HTML Assistant Pro interface, and Figure 21.2 shows the Microsoft FrontPage Editor window. HTML Assistant is a plain-text editor that adds an extensive button bar and many helpful menu options for dealing with common HTML tags. It also enables you to add your own tags as new extensions to HTML are introduced. FrontPage Editor, on the other hand, is a WYSIWYG editor that attempts to preview and automate as much of the editing process as possible.

Figure 21.1. HTML Assistant shows the HTML source text and offers an extensive toolbar.

Figure 21.2. MS FrontPage Editor is a what-you-see-is-sort-of-what-you-get style HTML editor.

The most obvious drawback to using any WYSIWYG editor is that you lose some control over the subtle details that would otherwise give your documents their unique, hand-crafted character. There is no substitute for working your way through each line of an HTML document, no matter how complex or simple your ideas might be. Although Web editors provide easier and speedier interfaces for many routine tasks, the changing nature of browsers like Internet Explorer (which allows for new HTML extensions on an almost monthly basis) defy any kind of automation, and hopefully encourage authors to explore the rich resources of HTML with more depth. Even FrontPage itself does not show or offer tools for creating some of the advanced features that Microsoft's own browser supports!

The HTML generated by WYSIWYG editors also tends to be quite difficult for humans to read (some editors even take out all extra spaces and line breaks on purpose to squeeze files to the absolute smallest possible size). This unreadability can become a major problem if you plan to maintain and update your Web site after the current version-du-jour has become obsolete.



NOTE

Some editors are a bit looser with official HTML syntax than others. The pickier editors might even refuse to open documents that use perfectly valid but non-standard extensions to HTML.


Regardless of these issues, the most important thing is that you work in a way and with tools that are comfortable for you. HTML Assistant is included on the enclosed CD-ROM, and FrontPage is available from Microsoft at http://www.microsoft.com/. You'll find detailed coverage of FrontPage in Part VI, "Building a Web Site with MS FrontPage." The chapters in this section focus on HTML as it would appear in a text-based editor such as HTML Assistant or when you choose View/HTML in the FrontPage Editor.

Putting Information on the Internet


Your Internet Service Provider (ISP) can give you the specific information you need about where on its computers you should place your HTML documents and advice on how best to get it there. Usually, you place your HTML documents in a directory called public_html or html_public or something similar. This public directory is unique among the other directories in your home directories in that any documents you place in it can and sometimes will automatically become readable to anyone on the World Wide Web. So be careful what you put there!

After you create your HTML documents, the easiest way to get them onto your ISP's server and into your public directory is to use some type of FTP client. FTP, as mentioned earlier, stands for File Transfer Protocol, and FTP clients such as WS_FTP or CuteFTP are programs that provide easy, point-and-click interfaces to that protocol.

Suppose I have several documents on my computer that I want to transfer to my public_html directory of my ISP server in Boston. Opening the WS_FTP program and clicking the Connect button brings up a simple dialog box into which I type my username, password, and the Internet address of my service provider, ftp.shore.net (see Figure 21.3).

Figure 21.3. To upload files to the Internet, you'll need an FTP program such as WS_FTP.

After I enter all this information and click OK, and assuming shore.net is its usual reliable self, I am connected to my home directory on shore.net. The interface is similar to that of the Windows File Manager or the Windows 95 Explorer, and I navigate from one directory to another by straightforward pointing and clicking. If I want to transfer files from my computer (which are listed on the left side of the screen) to shore.net, I just highlight the files I want to transfer and click the right-arrow button (see Figure 21.4). Most HTTP servers are set up so that any documents placed onto them are immediately made available to the entire World Wide Web.

Figure 21.4. To use FTP, highlight the files you want to send and press the transfer (arrow) button.



TIP

Some ISPs (Panix, for example) require that users manually change file permission settings, which control who is allowed to access individual files. You can use WS_FTP or CuteFTP to change permissions on an individual file. With WS_FTP, you select the file on the server and then click the right mouse button. A floating menu bar appears on-screen. Choose FTP commands, and then choose Site. In the Site dialog box, type chmod 755 YourFileNameHere.html. In English, this command means to change the mode of YourFileNameHere.html to 755. Changing the mode of this file to the numeric value of 755 tells the HTTP server that this file is authorized to be read by anyone on the World Wide Web. Your ISP might not require that you go through this whole business of changing file permissions and modes, so be sure to ask.



Getting Started in HTML


You have almost certainly encountered documents on the Web that made you ask "How did they do that?" Without question, the best and quickest way to learn HTML and to understand how it works is to frequently use Internet Explorer's View | Source command to view the HTML source code. Using this command might seem like cheating at first, but lay your guilt aside and know that learning by example is how virtually everyone first comes to understand and use HTML.

An HTML document is composed of individual HTML tags, also known as elements. (The terms tag and element are interchangeable for this discussion.) Each element serves functions at different levels of detail; HTML tags can determine the appearance of a single word, or they can be purely informational and simply state the content of the document itself.

HTML tags are contained within two-pronged brackets that you might know as "greater than" and "less than" symbols: < and >. Virtually all HTML tags are required to be "opened" and "closed." For example, you begin every document with an opening <HTML> tag, which indicates that this document is a Hypertext Markup Language document. A complementary closing </HTML> tag is also required. The closing tag always begins with a slash (/). Without the closing tag, any HTML tag is incomplete, and such incompleteness might cause Internet Explorer or other browsers to behave unpredictably. A few tags don't need to be closed. For example, the tag to insert a horizontal rule is just <HR>, with no closing </HR>. In general, however, every opening tag has a corresponding closing tag.

A Simple HTML Document


This section looks at a very simple HTML document (see Listing 21.1), and then looks at its document source so that you can understand why the document looks the way it does. This document contains the most basic elements of HTML, as well as a few more elegant markup tags. Figure 21.5 shows how it looks when loaded into a browser.

Figure 21.5. The HTML in Listing 21.1 when viewed with Internet Explorer.

Listing 21.1. A simple HTML document.




<HTML>



<HEAD>



<TITLE>Welcome to HTML 101</TITLE>



</HEAD>



<BODY>



<H1>How Do You Spell HTML?</H1>



This is your basic HTML document, and I am really chomping 



at the bit to put my life story on-line so I can make



<A HREF="http://www.microsoft.com/">a lot of money</A>.



</BODY>



</HTML>

The <HTML> Tag

The first element in an HTML document is the <HTML> tag. It tells browsers that access your page that this is a Hypertext Markup Language document, and that the information it contains is coded such that it should be interpreted through the filters of the HTML standards supported by that browser. (Different browsers support different levels of HTML.) The <HTML> tag and the <HEAD>, <TITLE>, and <BODY> tags discussed in the following sections are required in every HTML document. None of these tags necessarily affect the appearance of your document content, but without these elements, user agents might interpret your documents as ordinary text instead of HTML documents.

Remember that if you open a tag, you must close it at the appropriate place. Because the <HTML> element is the element that defines the document and the entire content of the document follows this element, you must place the closing </HTML> tag at the end of the document.

The <HEAD> and <TITLE> Tags

The <HEAD> section of a document is a collection of tags that contain information about the document itself. Within the <HEAD> element, the only required element is <TITLE>. After the <TITLE> tag, you enter what you want to call your document. In the example in Listing 21.1, the title is Welcome to HTML 101. This information, along with the name of the browser being used, appears in the title bar at the top of your browser (see Figure 21.5). Additionally, if someone adds your page to their Favorites list, the text within the <TITLE> and </TITLE> tags is what appears as the name of the page on that user's Favorites list.

The <BODY> Tag

The body of an HTML document is everything between the <BODY> and </BODY> tags. The body contains the content of your pages. For some pages, creating the body might involve taking pre-existing text and placing it into this new HTML document (or, for that matter, taking pre-existing text and surrounding it with HTML markup). For other pages, creating the body might mean typing new text straight into the HTML document. Either way, the <BODY> tag allows for the inclusion of more elaborate markup elements (known as attributes), such as colors and background images (which Chapter 22, "HTML 3 and Internet Explorer Extensions," explores).

Note that the document ends with the </BODY> and </HTML> closing tags. These closing tags are logically ordered according to the order in which their corresponding opening tags appeared. That is, because the opening <HTML> tag preceded the opening <BODY> tag, the closing </BODY> tag should appear before the closing </HTML> tag. Maintaining this basic structural symmetry throughout your HTML markup is central to ensuring that all your HTML documents work.

The essential template for any HTML document is as follows:




<HTML>



<HEAD><TITLE> </TITLE></HEAD>



<BODY>



</BODY>



</HTML>

All these tags might seem complex at first, but the basic sequence is the same for all HTML documents, so you have to learn it only once. You don't even need to remember it after that because you can always cut and paste the tags from your first page.



NOTE

Many HTML reference resources refer to a Document Type Declaration or <DTD> tag as being required for any document that is intended to be transmitted over a network. The purpose of the <DTD> statement is to alert the user agent that this information conforms to the document standards of the very specific document type announced in the document type declaration. The <DTD> typically looks something like the following:

<!DOCTYPE HTML PUBLIC "-//IETF//DTD HTML//EN//2.0">

As the purpose of the <DTD> is to clarify that this document conforms to a very specific HTML standard, its presence in HTML documents is not necessary because there is virtually no such thing as a pure HTML 2.0 or 3.x document. For a document to conform purely to the HTML 2.0 standard, it could contain only text and images, with nominal amounts of text formatting enhancements. For it to conform strictly to HTML 3.x (which is still not a defined standard), it would have to eliminate a variety of HTML 2 tags that are obsolete or replaced in HTML 3.



Heading Tags

The <H1> tag means Heading Level 1. There are six levels of headings (H1, H2, H3, H4, H5, and H6), and Heading Level 1 is typically the largest. Heading Level 1 is interpreted by almost all Web browsers as very large, boldfaced text, as shown in Figure 21.5. You find out more about heading tags a bit later in this chapter under "Basic Text Formatting."

Hypertext Links

<A HREF="http://www.microsoft.com/">a lot of money</A> is an anchor hypertext reference, more commonly known as a link. The text in the quotation marks is a uniform resource locator (URL), commonly just called an Internet address. The <A> and </A> tags are what make the Web a web; they provide hypertext hot links to other documents and enable you to jump around within a single document.

To insert a link to another document, use the HREF= attribute in the <A> tag to specify the filename of a local file or the address of any document or file on the Net. Any text that falls between the <A> and </A> tags is highlighted (in most browsers, colored and underlined). When readers click that text, they hop to the file you specified in the HREF= attribute. In Listing 21.1, clicking on the words a lot of money takes you right to the Microsoft Corporation Home Page.

You can also use the <A> tag to jump to other parts of an HTML document by creating a named reference point, as in the following:




<A NAME="introduction"></A>Once upon a time...

You can then jump to the introduction anchor by creating a link from some other part of the document. Put a # character before the anchor name in the HREF= attribute as follows:




<A HREF="#introduction">Click here</A> to go to the introduction.

Plan to Be Mangled


An individual HTML document looks different when other people view it. An HTML document's appearance depends not only on the browser users use, but also on the way they configure their monitors, modems, and browsers. Most browsers, including Internet Explorer, offer users elaborate personal preferences so that the appearance of documents can be customized in ways Web authors might not have predicted.

Figures 21.6 and 21.7 show the same HTML document as it looks in Internet Explorer version 3.0 with different user option settings. Note that the font size, window size, and background color setting (which causes the smudges in Figure 21.6) can have a dramatic effect on how text wraps and how a page appears.

Figure 21.6. Part of a fairly simple Web page displayed in the Internet Explorer browser.

Figure 21.7. The Web page in Figure 21.6 viewed in the same version of Internet Explorer, but with different user options in effect.

Of course, the page looks even more different when you use different software to access the Web, as demonstrated in Figure 21.8. Other factors, such as the resolution and number of colors that a user's video hardware can display, can also have a profound effect on how a page looks.

Figure 21.8. The same Web page as in the two previous figures viewed in the DOS-based Lynx browser.

The discrepancies between Figures 21.6, 21.7, and 21.8 are a source of frustration that's quite in line with the spirit of HTML, which is intended to be a content-oriented (rather than appearance-oriented) standard, but they do imply two things for people who carefully design and lay out colorful graphics on their Web pages. First, always check to see how your pages look in various common browsers with various preferences settings and window sizes before you post them on the Net. Second, don't get too hung up on fine-tuning the exact appearance of your page in a particular browser configuration. Instead, focus on general organization and strong content that is compelling no matter how the reader looks at it.

By including support for elements from the proposed HTML 3.2 specification, Internet Explorer allows for considerably more splashy document markup than some other browsers. Such elements as tables, colors, dynamic documents, frames, and numerous other text placement and formatting tags (covered in Chapter 22) are part of HTML 3. For now, virtually none of the widely used Web browsers interpret HTML strictly according to either HTML 2 or 3. Most conform to a healthy mixture of HTML 2 and HTML 3 elements.

Basic Text Formatting


Suppose I wanted to make a Web page out of the following text:

All species surveyed noted the lack of olfactory media on the Web

To tell the Web browser which part of the text is a heading, I insert tags at the beginning and end of each heading. The <H1> and </H1> tags begin and end the largest size heading, as shown in the following example:




<H1>Bowser and Bootsie Agree!</H1>

I can also use <H2>, <H3>, <H4>, <H5>, and <H6> to make smaller subheadings. (Although six levels of headings are available, try to stick to only three levels; the differences between <H3>, <H4>, <H5>, and <H6> are rather subtle.)




<H1>Bowser and Bootsie Agree!</H1>



<H2>Internet Explorer Voted Best Cross-Species Browser</H2>

Note that Internet Explorer (and almost all other browsers) inserts a blank line after the </H1> tag. Ordinarily, if you wanted to include a blank line in an HTML document, you would use a <BR> or <P> tag. Because heading elements are intended to contain headers that precede or otherwise demarcate sections of text, a line return is implied and is therefore automatically inserted.

You must enclose the entire body of an HTML document between the <HTML>, <BODY>, <HEAD> and <TITLE> tags discussed earlier in this chapter. You can use the <A> tag, also mentioned earlier, to add hypertext links to other Web pages. You might, for example, insert a link to the Non-Human Society home page:




<A HREF="http://www.nonhuman.org/">Non-Human Society home page</A>

To transform the rest of the text into a respectable Web page, you would need a number of other text formatting tags. A normal paragraph of text can begin with a <P> tag and end with </P>, but the HTML conventions enable you to take a shortcut and simply put a single <P> tag at the beginning or end of each paragraph. The <P> tag usually is interpreted to insert two line breaks, skipping a line between paragraphs. To insert a line break without skipping a line, use the <BR> tag.

You can use <I> and </I> to begin and end a section of italicized text. Similarly, <B> and </B> begin and end boldface text. Some HTML purists prefer to use the <EM> and </EM> tags for emphasis and the <STRONG> and </STRONG> tags for strong emphasis, leaving the decision of exactly how to render the emphasis up to the browser. Although <EM> and <STRONG> might be more in line with the philosophy of markup languages, almost all Web page authors opt for the control and simplicity of <I> and <B>.

You can also indicate numbered or bulleted lists. Numbered lists are called ordered lists; they begin with the <OL> tag and end with </OL>. The browser automatically adds the numbers (1, 2, 3, and so on). Bulleted lists are called unordered lists; they begin with <UL> and end with </UL>. Each new item on the list begins with an <LI> tag, and you can put an optional </LI> tag at the end of each line if you want to.

Putting all this information together, a simple but complete hypertext document made by marking up the text given previously might look like the passage in Listing 21.2. (Figure 21.9 shows the following HTML as it appears when viewed with Internet Explorer.)

Figure 21.9. The HTML in Listing 21.2 as interpreted by Internet Explorer.

Listing 21.2. A sample hypertext document.




<HTML>



<HEAD><TITLE>Canine/Feline Survey Report</TITLE></HEAD>



<BODY>



<H1>Bowser and Bootsie Agree!</H1>



<H2>Internet Explorer Voted Best Cross-Species Browser</H2>



There seems to be <B>one</B> thing cats and dogs agree on these days:



In a recent survey of impounded animals with Internet access,



both canine and feline respondents voted <I>Microsoft Internet Explorer</I>



as their favorite Web browser.<P>



The study, which can be viewed in detail at the



<A HREF="http://www.nonhuman.org/">Non-Human Society home page</A>,



was the first systematic census of non-primate Web users. 



Other results reported include:<BR>



<ul>



<LI>Cats are more Internet aware, browsing on average 50% more than dogs



<LI>Dogs were 80% more likely to turn automatic image loading off



<LI>All species surveyed noted the lack of olfactory media on the Web



</UL>



</BODY>



</HTML>

Table 21.1 summarizes the most commonly used tags. As you can see, they are easy to learn and use.

Table 21.1. The most common HTML tags.

Usage Opening tag Closing tag
Entire document <HTML> </HTML>
Document header <HEAD> </HEAD>
Title (within header) <TITLE> </TITLE>
Document body <BODY> </BODY>
Top-level heading <H1> </H1>
2nd-level heading <H2> </H2>
3rd-level heading <H3> </H3>
Italic text <I> </I>
Bold text <B> </B>
Monospaced text <TT> </TT>
New paragraph <P> (</P> is optional)
Line break (within a paragraph) <BR>
Horizontal rule <HR>
Ordered (Numbered) list <OL> </OL>
Unordered (Bulleted) list <UL> </UL>
New line in list <LI> (</LI> is optional)
Image <IMG>
Anchor/link <A> </A>

Inline Images


The <IMG> tag tells the browser to insert an image. You must specify where to find the image file and how to place the next line of text in relation to the image. The image location, called the source, can be a filename on the same computer as the HTML text file or an URL pointing to a file on another computer. Usually, you put the images in the same directory as the HTML pages, so you can specify a filename with the SRC= attribute, as follows:




<IMG SRC="catdog.gif">


TIP

Internet Explorer supports inclusion of GIF, JPG, and BMP images only—and most other browsers can view only GIF and JPG. Other image formats can be viewed using external helper applications and plug-ins, but you should convert all images to GIF and JPG for inclusion in the main body of a Web page. (The Paint Shop Pro software on the CD-ROM can convert almost any graphics format to GIF or JPG. See Chapter 24, "Graphics and Multimedia for Internet Explorer," for more details.)


The next line of text after an image is assumed to be a caption and is placed immediately to the right of the image. Using the ALIGN= attribute, you can specify whether the caption should be aligned with the TOP, MIDDLE, or BOTTOM of the image, as in the following example:




<IMG SRC="catdog.gif" ALIGN="TOP">A public service announcement.

You should enclose the word TOP, MIDDLE, or BOTTOM in quotes, although most browsers enable you to cheat and leave the quotes off. If you don't want the next line of text to be placed next to the image, put a <P> tag after the <IMG> tag, and leave out the ALIGN= attribute. Internet Explorer also supports several other nonstandard options for the ALIGN attribute, which are discussed in Chapter 23, "HTML 3.0 and Internet Explorer Extensions."

Figure 21.10 shows the page from Figure 21.9 with the addition of an inline graphic image.

Figure 21.10. Graphics add interest and color to Web pages and are easy to incorporate.

If you include an image within a link, a colored border appears around the image. When the user clicks any part of the image, he or she jumps to the address in the link's HREF. For example, if you wanted to allow readers to jump to a document named thepound.htm by clicking an icon named catdog.gif or an accompanying caption that read "Pet Us!," you would write the following:




<A HREF="thepound.htm">



<IMG SRC="catdog.gif">Pet Us!</A>


TIP

Remember that many people will not see any of the graphics you put on your Web pages. Even those using graphics-capable browsers often surf with graphics downloading turned off to reduce download times over a slow modem connection.

HTML gives you a way to send a special message to readers who don't see your graphics. You can give each image an ALT= attribute that displays the text you specify whenever the image itself can't be shown, for example:

<IMG SRC="triangle.gif" ALT="WARNING:">
This page is radioactive on some monitors.

In this example, if the image file triangle.gif can't be displayed, most browsers would display the word WARNING: instead.

If an image is also a clickable link, Internet Explorer will display the ALT text in a small box over the image whenever a user's mouse passes over the image. This allows you to give people a brief text "preview" or "hint" about what they'll get if they click over an image.



Forms


Web forms enable you, the Web page publisher, to receive feedback, orders, or other information from the readers of your Web pages. All major browsers support forms, and their syntax is included in the HTML 2.0 standard specification. The rest of this chapter shows you how to create your own forms and the basics of how to handle form submissions. You'll find out about many more advanced uses for forms in Part VII, "Web Scripting and ActiveX."

Creating forms for use on the World Wide Web is not complicated or difficult, as many newcomers to Web publishing assume it must be. Even a beginner can usually create a useful form on the first try. Figure 21.11 shows all the form types as a reminder and reference.

Figure 21.11. A whimsical form, demonstrating all the HTML input elements.



TIP

Notice that most of the text in Figure 21.11 is monospaced. Monospaced text makes it easy to line up a form input box with the box above or below it and makes your forms look neater. To use monospaced text throughout a form, enclose the entire form between <PRE> and </PRE> tags. Using these tags also relieves you from having to put <BR> at the end of every line because the <PRE> tag puts a line break on the page at every line break in the HTML document.



The <FORM> Tag


Every form must begin with a <FORM> tag, which can be located anywhere in the body of the HTML document as long as it occurs before the first INPUT tag. The FORM tag normally has two attributes, METHOD and ACTION:




<FORM METHOD="POST" ACTION="/cgi/generic">

The body of the form follows this line and ends with the following line:




</FORM>

Nowadays, the METHOD is almost always "POST", which means to send the form entry results as a document. (In some special situations, you might need to use METHOD="GET", which submits the results as part of the URL header instead. For example, "GET" is sometimes used when submitting queries to search engines from a Web form. (If you're not yet an expert on forms, just use "POST" unless someone tells you to do otherwise.)

The ACTION attribute points to the program or script on the server computer that will process the information that a user enters on a form. If your Web site is hosted by an Internet Service Provider (ISP), it will usually have a preinstalled library of scripts to handle the most common sorts of actions, such as e-mailing the form data to you and generating a confirmation response page for the user. Your ISP can tell you how to use these scripts. If you are running your own server, consult your server software documentation to find out what scripts (if any) come with it. You can also try writing your own scripts in any language supported on the server. See Chapter 31, "CGI Server-Side Scripting," for help.

Text Input


To ask the user for input, you use the <INPUT> tag. This tag must fall between the <FORM> and </FORM> tags, but it can be anywhere on the page in relation to text, images, and other HTML tags. For example, you would ask the user's name in the following manner:




What's your first name? 



<INPUT TYPE="text" SIZE=20 MAXLENGTH=30 NAME="firstname">



What's your last name? 



<INPUT TYPE="text" SIZE=20 MAXLENGTH=30 NAME="lastname">

The TYPE attribute indicates what type of form element to display, a simple one-line text entry box in this case. (Each element type is discussed individually in the following sections.)

The SIZE attribute indicates approximately how many characters wide the text input box should be. If you are using a proportionally spaced font, the width of the input will vary depending on what the user enters. If the input is too long to fit in the box, Netscape (and most other browsers) will automatically scroll the text to the left.

MAXLENGTH determines the number of characters that the user is allowed to type into the text box. If the user tries to type beyond the specified length, the extra characters do not appear. You might specify a length that is longer, shorter, or the same as the physical size of the text box. SIZE and MAXLENGTH are used only for TYPE="text" because other input types (check boxes, radio buttons, and so on) have a fixed size.

No matter what type an input element is, however, you must give a name to the data it gathers. You can use any name you like for each input item, as long as each one on the form is different. When the form is sent to the server script specified in the FORM ACTION attribute, each data item is identified by name.

For example, if the user entered Jane and Doe in the text box defined previously, the server script would receive a document including the following:




firstname='Jane'



lastname='Doe'

If the script mails the raw data from the form submission to you directly, you would see these two lines in an e-mail message from the server. Or the script might be set up to format the input into a more readable format before reporting it to you.



TIP

If you want the user to enter text without it being displayed on the screen, you can use INPUT TYPE="password" instead of INPUT TYPE="text". Asterisks (***) are then displayed in place of the text the user types. The SIZE, MAXLENGTH, and NAME attributes work exactly the same for TYPE="password" as for TYPE="text".



Check Boxes


The simplest input type is a check box, which appears as a small square that the user can select or deselect by clicking. A check box doesn't take any attributes other than NAME:




<INPUT TYPE="checkbox" NAME="widget" VALUE="yes"> Standard Widget



<INPUT TYPE="checkbox" NAME="sdwidget" VALUE="yes"> Super-Deluxe Widget

Selected check boxes appear in the form result sent to the server script as follows:




sdwidget='yes'

Blank (deselected) check boxes do not appear in the form output result at all. If you don't specify a VALUE attribute, the default VALUE of "yes" is used.



TIP

You can use more than one check box with the same name, but different values, as in the following code:

<INPUT TYPE="checkbox" NAME="pet" VALUE="dog"> Dog
<INPUT TYPE="checkbox" NAME="pet" VALUE="cat"> Cat
<INPUT TYPE="checkbox" NAME="pet" VALUE="iguana"> Iguana

If the user checked both Cat and Iguana, the submission result would include the following:

pet='cat'
pet='iguana'



Radio Buttons


Radio buttons, where only one choice can be selected at a time, are almost as simple to implement as check boxes. Just use TYPE="radio" and give each of the options its own INPUT tag, as in the following code:




<INPUT TYPE="radio" NAME="card" VALUE="v" CHECKED> Visa



<INPUT TYPE="radio" NAME="card" VALUE="m"> MasterCard

The VALUE can be any name or code you choose. If you include the CHECKED attribute, that button will be selected by default. (No more than one button with the same name can be checked.)

If the user selected MasterCard from the preceding radio button set, the following would be included in the form submission to the server script:




card='m'

If the user didn't change the default CHECKED selection, card='v' would be sent instead.

Selection Lists


Both scrolling lists and pull-down pick lists are created with the <SELECT> tag. You use this tag together with the <OPTION> tag:




<SELECT NAME="extras" SIZE=3 MULTIPLE>



<OPTION SELECTED> Electric windows



<OPTION> AM/FM Radio



<OPTION> Turbocharger



</SELECT>

No HTML tags other than <OPTION> should appear between the <SELECT> and </SELECT> tags.

Unlike the text input type, the SIZE attribute here determines how many items show at once on the selection list. If SIZE=2 had been used in the preceding code, only the first two options would be visible, and a scrollbar would appear next to the list so the user could scroll down to see the third option.

Including the MULTIPLE attribute enables users to select more than one option at a time, and the SELECTED attribute makes an option selected by default. The actual text accompanying selected options is returned when the form is submitted. If the user selected Electric windows and Turbocharger, for example, the form results would include the following lines:




extras='Electric windows'



extras='Turbocharger'


TIP

If you leave out the SIZE attribute or specify SIZE=1, the list will create a pull-down pick list. Pick lists cannot allow multiple choices; they are logically equivalent to a group of radio buttons. For example, the following is another way to choose between credit card types:

<SELECT NAME="card">
<OPTION> Visa
<OPTION> Mastercard
</SELECT>



Text Areas


The <INPUT TYPE="text"> attribute mentioned earlier enables the user to enter only a single line of text. When you want to allow multiple lines of text in a single input item, use the <TEXTAREA> and </TEXTAREA> tags instead. Any text you include between these two tags will be displayed as the default entry. Here's an example:




<TEXTAREA NAME="comments" ROWS=4 COLS=20>



Please send more information.



</TEXTAREA>

As you probably guessed, the ROWS and COLS attributes control the number of rows and columns of text that fit in the input box. Text area boxes do have a scrollbar, however, so the user can enter more text than fits in the display area.



NOTE

Some older browsers do not support the placement of default text within the text area. In these browsers, the text might appear outside the text input box.


Internet Explorer now supports the following new TEXTAREA WRAP attributes to control how text wraps to the next line when the user reaches the end of a line.


Submit


Every form must include a button that submits the form data to the server. You can put any label you like on this button with the VALUE attribute, as in the following:




<INPUT TYPE="submit" VALUE="Place My Order Now!">

A grey button will be sized to fit the label you put in the VALUE attribute. When the user clicks it, all data items on the form are sent to the program or script specified in the FORM ACTION attribute.

Normally, this program or script generates some sort of reply page and sends it back to be displayed for the user. If no such page is generated, the form remains visible, however.

You may also optionally include a button that clears all entries on the form so users can start over again if they change their minds or make mistakes. Use the following line:




<INPUT TYPE="reset" VALUE="Clear This Form and Start Over">


TIP

If you want to send certain data items to the server script that processes a form, but you don't want the user to see them, you can use the INPUT TYPE="hidden" attribute. This attribute has no effect on the display at all; it just adds any name and value you specify to the form results when they are submitted.

You might use this attribute to tell a script where to e-mail the form results. For example, this line

<INPUT TYPE="hidden" NAME="mail_to" VALUE="me@somewhere.com">

would add the following line to the form output:

mail_to='me@somewhere.com'

For this attribute to have any effect, someone must create a script or program to read this line and do something about it. You might also use hidden items to indicate which of many similar forms a particular result came from.



Uploading Files


A new and seldom used addition to forms is the ability for users to upload entire files from their hard drive as part of a Web form submission. Uploading is done with INPUT TYPE="file", which displays a text box where the user can enter the path and name of the file to upload.

File uploads should be the only input item within a form because they require that you specify the FORM ENCTYPE="multipart/form-data" attribute. This attribute instructs the browser to send the form results using standard MIME encoding, which is necessary in order to allow non-text data files to be included. An example form would be as follows:




<FORM ENCTYPE="multipart/form-data" ACTION="http://formaction" METHOD=POST>



Please enter the name of the file you wish to send: 



<INPUT NAME="userfile" TYPE="file">



<INPUT TYPE="submit" VALUE="Send File">



</FORM>

What's Next?


This chapter has taken you through a crash course in HTML. The following chapters discuss a number of more sophisticated aspects of HTML, including forms, frames, fonts, and fancy formatting.

Previous Page Page Top TOC Next Page