This chapter contains reference information about encodings, emulations, and configuring SlickEdit Core.
Encodings are used to convert a file to either SBCS/DBCS for the active code page or Unicode (more specifically UTF-8) data. By default, XML and Unicode files with signatures (UTF-8, UTF-16 and UTF-32) files are automatically loaded as Unicode UTF-8 data, while other more common program source files like .c, .java, and .cs source files are loaded as SBCS/DBCS active code page data.
All file data can be configured to Unicode UTF-8 data, but this would cause some problems. Loading files containing SBCS/DBCS data would take significantly longer, slowing down parsing by Context Tagging® and any other multi-file operations. In addition, Unicode editors cannot support all the features supported by SBCS/DBCS editors due to font limitations. For more information, see Unicode Limitations.
To provide better support for editing Unicode and non-Unicode files, two modes of editing exist: Unicode and SBCS/DBCS mode. Files that contain Unicode, XML, or code page data not compatible with the active code page should be opened as Unicode files.
The following are non-Unicode encodings and put the editor in SBCS/DBCS editing mode: Default, Text, SBCS/DBCS mode, Binary, SBCS/DBCS mode, and EBCDIC, SBCS/DBCS mode. In addition, the Auto Unicode, Auto Unicode2, Auto EBCDIC and Unicode, and Auto EBCDIC and Unicode2 encodings put the editor into SBCS/DBCS editing mode when the file is determined not to be Unicode. All other encodings put the editor in Unicode mode and require that the file data be converted to UTF-8.
There are many encodings available, including:
Auto XML - This encoding specifies that the file encoding be determined based on XML standards and that the file be loaded as Unicode data. The encoding is determined based on the encoding specified by the ?xml tag. If the encoding is not specified by the ?xml, the file data is assumed to be UTF-8 data which is consistent with XML standards. We applied some modifications to the standard XML encoding determination to allow for some user error. If the file has a standard Unicode signature, the Unicode signature is assumed to be correct and the encoding defined by the ?xml tag is ignored.
Auto Unicode - When this encoding is chosen and the file has a standard Unicode signature, the file is loaded as Unicode data. Otherwise the file is loaded as SBCS/DBCS data.
Auto Unicode2 - When this encoding is chosen and the file has a standard Unicode signature or looks like a Unicode file, the file is loaded as Unicode data. Otherwise the file is loaded as SBCS/DBCS data. This option is NOT fool-proof and may give incorrect results.
Auto EBCDIC - When this encoding is chosen and the file looks like an EBCDIC file, the file is loaded as Unicode data. Otherwise, the file is loaded as SBCS/DBCS data. This option is NOT fool-proof and may give incorrect results. The option does attempt to support binary EBCDIC files.
Auto EBCDIC and Unicode2 - This encoding is a combination of the Auto EBCDIC and Auto Unicode2 encodings described above.
To use encodings in SlickEdit® Core, Unicode support is required (OEMs typically turn this feature off). Unicode is supported for the following list of features:
All Context Tagging® features.
Color Coding.
Level 1 regular expressions as defined by the Unicode consortium.
Multi-file search and replace.
Support for many encodings including UTF-8, UTF-16, UTF-32, and many code pages. Automatic encoding recognition for XML files. Configure encoding recognition per extension or globally. Optionally store signatures and specify little endian or big endian. Use the Save As or Write Selection dialog to convert data to a particular file encoding.
Support for converting Unicode to UNC data and visa versa. Supported UCN formats include \xHHHH, \x{HHHH}, \uHHHH, &xHHHH;, and &xDDDD;. This is useful for specifying Unicode character strings in SBCS/DBCS active code page source files. See Converting Unicode to UCN.
Multiple clipboards.
Sorting.
3-Way Merge.
Support for composite and surrogate characters.
Support for storing up to 31-bit Unicode characters.
SmartPaste®.
Syntax Expansion and Syntax Indenting.
Code beautifiers.
Support for almost all of SlickEdit Core's SBCS/DBCS active code page features.
By default, XML and Unicode files with signatures (UTF-8, UTF-16 and UTF-32) files are automatically loaded as Unicode. If you have Unicode files that are not XML and do not have signatures, configure default options to get the best recognition possible. This is important because some features such as drag/drop files and DIFFzilla® do not prompt you for the file encoding.
Each extension may have its own encoding specification. If the extension-specific encoding is set to Default, then the global setting defined at Window → SlickEdit Preferences → Languages → File Extension Manager is used. Both the extension-specific and global setting are overridden if you previously specified an encoding in the Open dialog. The encoding used to override default encoding settings is recorded. The setting is then reused the next time you open the same file. This provides you with per-file encoding support.
If you have non-XML UTF-16 files that have signatures, then try selecting Auto Unicode2 as an extension-specific or global encoding. Since there is no option for recognizing UTF-8 or UTF-32 files (other than Auto XML) by looking at the file contents, you will either need to set an extension-specific encoding, or specify the encoding in the Open dialog the first time you open the file.
Some compilers (such as Visual C++) let you specify the code page in the source file (in fact, more than one code page can be used in the file). This is not supported, so the assumption is that the file is SBCS/DBCS active code page data.
To open a Unicode file, complete the following steps:
Use the Open dialog (File → Open).
Specify the encoding if necessary.
Press Enter.
Unicode data is stored as UTF-8 and not UTF-16. Since the Windows Win32 calls are used to implement some Unicode features there are some issues. By default, Windows does not support surrogates. You must use the regedit program to turn on surrogate support.
To turn on surrogate support, run the regedit program and go to the following key location:
HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows NT\CurrentVersion\LanguagePack
Set the value for SURROGATE to 0x00000002.
Casing features (uppercase, lowercase, ignore case) do not support surrogates. Windows is used for casing support and Windows casing features do not support surrogates.
You can convert a selection from Unicode to UCN or vice-versa. SlickEdit® Core conversion features are located on the Edit → Other menu. The Edit → Other → Unicode to UCN conversion feature is most useful for specifying Unicode character strings in SBCS/DBCS active code page source files. For example, here are the steps to store some UCN in a Java source file:
Open the Unicode file containing the Unicode characters or create a new Unicode file and enter the characters you want to convert.
Select the Unicode characters you want to convert.
Execute the Edit → Other → Copy Unicode As → Java/C# (UTF-16 \uHHHH) menu item.
Open the Java source file and paste (Edit → Paste) the UCN data into the file.
The following is a list of Unicode limitations:
Bold and italics color-coding is not supported. Support for this will be added in a future version.
Tab character operations are not fully supported. Tab display, the Expand tabs to spaces save option (Window → SlickEdit Preferences → File Options → Save), and save with tabs (save +t) only work correctly if all the characters are below 128. The Expand tabs to spaces load option (Window → SlickEdit Preferences → File Options → Load) is ignored.
Column selections do not fully support Unicode. If all the characters are below 128 and the font is fixed then it works. Support for this will be added in a future version.
Word Wrap does not fully support Unicode. If all the characters are below 128 and the font is fixed, then it works. Support for this will be added in a future version.
The Unicode line end character 0x2048 is not supported.
Hex editing is not supported. The current character (Composite character) is displayed on the status line. Also, use the Open dialog with the Binary, SBCS/DBCS mode encoding to view a Unicode file in hexadecimal.
Casing features (uppercase, lowercase, ignore case) do not support surrogates. Windows is relied upon for casing support, and Windows casing features do not support surrogates. See Surrogate Support.
Vertical line column (Window → SlickEdit Preferences → Appearance → General) is not supported.
Truncation line length is not supported.
Record width on the File Open dialog is not supported.
DDE is not supported. Unicode DDE does not work with Internet Explorer or Netscape®. You can view files with Unicode data in Internet Explorer; however, this feature will fail if the file name contains characters not in the active code page.
Version control supports files containing Unicode data but does not support file names that contain characters not in the active code page.
Special character display is not supported for Unicode buffers.
The grew program does not support Unicode and can only be used on SBCS/DBCS active code page text.
If you load the same source file in Unicode and SBCS/DBCS mode, the Context Tagging® database will have incorrect seek positions. It is important to use the default load options and to always load source files in the same encoding so that the Context Tagging seek positions match the editor seek positions.
The install (setup.exe), unionist (uninstall.exe), and update (update.exe) programs are not Unicode applications so the installation directory must contain characters in the active code page.
Native Unicode and SBCS/DBCS editing modes are supported. When you edit a SBCS/DBCS (active code page) file such as a .c, .h, or .java file, the data is loaded as SBCS/DBCS data and is not converted to Unicode. When you edit a Unicode file, such as an XML file, the data is converted to UTF-8 that is one of the standard formats for supporting Unicode files. There are several advantages to this implementation:
Since almost all source files for programming are stored as SBCS/DBCS, loading these files is significantly faster. This is very important to our customers who expect superior performance from SlickEdit® Core.
Unicode editing modes cannot support all the features you were used to when editing SBCS/DBCS files (see Unicode Limitations).
Macros can be written once to support both editing modes. This was very important to us because we wanted to reduce development time.
Since Unicode is stored as UTF-8, only one set of binaries is required. Most products that support SBCS/DBCS and Unicode (UTF-16), use preprocessing. This requires two sets of binaries.