Regular expressions are patterns of text used to match and manipulate strings in your code. These patterns are expressed with combinations of characters defined by the regular expression syntax being used. A regular expression is sometimes referred to as a "regex".
Use regular expressions in your search and replace operations when you find normal search/replace too limiting. For example, with regular expressions, you can:
Find quoted strings.
Find blank lines.
Find words starting at the beginning of lines.
Find two words separated by any number of spaces or other text.
SlickEdit® Core supports five types of regular expression syntax:
SlickEdit Core also provides a Regex Evaluator that you can use to interactively create, save, and re-use tests of regular expressions. See The Regex Evaluator for more information.
See Regular Expression Syntax for charts of the expressions in syntax. Unicode regular expression categories and character blocks are also supported. See Unicode Categories and Character Blocks for more information.
This documentation is not meant to be an exhaustive resource on regular expressions. Rather, we will present basic information, syntax charts, and examples. For novice users, there are many books and Web sites that go into more detail about this topic.
While regular expressions in SlickEdit Core primarily match the syntax for that language, there are some differences between our implementation and those elsewhere.
All search and replace commands, the Find and Replace view, and incremental search support regular expressions. For search and replace commands and the view, you can specify the regular expression syntax to use through specific options. A global option is available to specify the default syntax to use when you invoke these features or when you use incremental search.
For example:
Search and replace commands - When using the search commands / (slash) and find, or the replace commands c and replace, you can use the following options to specify regular expression syntax:
Use B to interpret the search string as Brief regular expression.
Use the L option to interpret the string as a Perl regular expression.
Use R to interpret the string as a SlickEdit regular expression.
Use the U option to interpret the string as a UNIX regular expression.
Find and Replace view - When using the view, select Use in the Search options box, and then pick the syntax to use from the drop-down list.
Incremental search - When using incremental search, press Ctrl+T to toggle regular expression searching on and off. The syntax that will be used is based on the global syntax setting.
To set the global option, from the main menu, click Window → SlickEdit Preferences, expand Editing and select Search. Set the Regular expression option to True and select the syntax you want to use from the Expression type drop-down list.
If you are using tagged expressions or regular expressions to perform a search and replace, it is important to understand the difference between the minimal and maximal operators.
Take, for example, a line of text which contains a DOS file name: \path1\path2\path3\name.ext.
Based on the syntax, the following regular expressions match the string \path1\:
The following regular expressions, which use the maximal operator, match the string \path\path2\path3\:
As a rule of thumb, the following minimal matching operators are generally used after a less-specific regular expression such as ? in Brief/SlickEdit or .in UNIX:
Use the maximal matching operators after a regular expression which matches something more specific. For example, to search for a string of digits and prefix each matched string with the character $, specify the following expressions:
|
Syntax |
Expression |
|---|---|
|
Brief |
Search for: {[0-9]\:+} Replace with:$\0 |
|
Perl and UNIX |
Search for:([0-9]+) Replace with:$\1 |
|
SlickEdit |
Search for: {[0-9]#} Replace with: $#0 |
If the minimal matching operator (+ in Brief/SlickEdit syntax, +? in Perl/UNIX) was used in the search string instead of the maximal matching operator (\:@ in Brief, + in Perl/UNIX, # in SlickEdit Core), the above search and replace would prefix each digit in the entire file with a $ character.
When you use regular expressions to search for a string, you will often want the replace string to depend on what was found. Use tagged expressions to insert parts of what is found into the replace string. Tagged expressions are denoted based on the syntax:
Brief syntax - Use { } (curly braces) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) followed by a tagged expression number 0-9. Count the { (left braces)in the search string to determine a tagged expression number. The first tagged expression is \0.
UNIX syntax - Use ( ) (parentheses) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) followed by a tag group number 1-9. Count the ( (left parenthesis) in the search string to determine a tagged expression number. The first tagged expression is \1 and the last is \0.
Perl syntax - Use ( ) (parentheses) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) or $ (dollar sign) followed by a tag group number 1-9. Count the ( (left parenthesis) in the search string to determine a tagged expression number. The first tagged expression is \1 and the last is \0.
SlickEdit® Core syntax - Use { } (curly braces) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a # (pound sign) followed by a tagged expression number 0-9. Count the { (left braces) in the search string to determine a tagged expression number. The first tagged expression is #0.
Examples of Tagged Expressions
The expressions in the table below replace occurrences of "if" and "while" with "xify" and "xwhiley." Unmatched groups are null. Note that the \1 in Brief syntax, \2 in Perl/UNIX, and #1 in SlickEdit syntax are replaced with null.
The expressions in the table below reverse text on lines containing a comma. Lines with "abc,def" will be changed to "def,abc". Notice that the Perl/UNIX regular expression search string uses a *? minimal matching operator, so the comma actually matches the first comma in the line and not the last.
When using regular expressions, some characters have a different meaning when used in the replace string, depending on the syntax:
Brief - The backslash character (\) in the replace string has the same meaning as in the search string except that \c and \:char are not supported.
UNIX - A backslash in the replace string has the same meaning as in the search string except that \c and \:char are not supported.
Perl - A backslash in the replace string has the same meaning as in the search string except that \c and \:char are not supported. A dollar sign ($) must be escaped (\$) when replacing a literal $.
SlickEdit® Core - The pound sign character (#) and backslash (\) have special meaning in the replace string. A backslash in the replace string has the same meaning as in the search string except that \c, :char, and \gd are not supported.
See the Regular Expression Syntax tables for a list of options for these characters. See Using Tagged Expressions for information on specifying tagged expressions in the replace string.
When used in a replace operation, the expressions in the following table can be used to modify the character casing of matched expressions. These work in Brief, Perl, SlickEdit, and UNIX syntaxes.
The table below contains some examples of replace operations using regular expressions.
|
Operation |
Expression |
|---|---|
|
Search for occurrences of the string "hat" that occur at the end of a line and replace it with "cat". |
In all syntaxes: Search for: hat$ Replace with: cat |
|
Delete blank lines. |
Brief: Search for: <\n Replace with: (leave blank) Perl/SlickEdit/UNIX: Search for: ^\n Replace with: (leave blank) |
|
Replace occurrences of two consecutive blank lines with one. |
Brief: Search for: <\n\n Replace with: \n Perl/SlickEdit/UNIX: Search for: ^\n\n Replace with:\n |
|
Search for lines containing "a" and replace the "a" with a formfeed character. |
Brief: Search for: <a+$ Replace with:\d12 Perl/UNIX: Search for:^a+$ Replace with:\d12 SlickEdit: Search for:^a+$ Replace with:\12 |
|
Select occurrences of "Title:" at the beginning of a line and capitalize the text following "Title:". |
Brief: Search for: <Title: {\:*} Replace with: Title: \U\0 Perl/UNIX: Search for: ^Title: (.*) Replace with: Title: \U\1 SlickEdit: Search for: ^Title\: {?@} Replace with: Title: \U#0 |
Regular expressions are used to express text patterns for searching. The Regex Evaluator provides the capability to interactively create, save, and re-use tests of regular expressions.
To access the Regex Evaluator, click Tools → Regex Evaluator (or use the activate_regex_evaluator command). Like other views in SlickEdit® Core, this view is dockable. Docking options can be accessed by right-clicking on the view's title bar.
Type some samples of the text you are trying to match in the top portion of the view labeled Test Cases. Enter your regular expression pattern in the bottom field. The Regex Evaluator will highlight matched portions of your sample text and identify groups.

Type your test cases in the Test Cases text box. These test cases will be evaluated as you type your regular expression in the bottom field. A wavy underline will indicate the ranges of text that match the entire expression. Matches are also marked with a yellow arrow that appears in the gutter to the left of the test case. You can hover your mouse on this arrow to see a tool tip which displays the matched expression details. When groups (tagged expressions) are used in your regular expression pattern, the groups will be boxed and highlighted in yellow in the Test Cases section.
Enter the regular expression to test in the text field. Use the radio buttons to select the expression syntax that you wish to use: UNIX, SlickEdit® Core, Brief, or Perl. Click the arrow to the right of the regular expression field to pick from a menu of common syntax and operators.
The following options and buttons are available on the Regex Evaluator view:
Multiline mode - If Multiline mode is selected, rather than searching through the test cases line-by-line, regular expressions will be searched on all lines at once. This is useful for test cases that wrap to the next line. This works just as if you had entered \om on the SlickEdit® Core command line.
Case sensitive - If Case sensitive is selected, the regular expression search will be case sensitive. This option is on by default.
New expression button - To clear the view of all entries in order to start a new evaluation, click the button at the top of the view labeled New expression.
Open a saved expression button - To open an expression that you have already saved, click the folder button at the top of the view labeled Open a saved expression.
Save the current expression button - To save the current expression, click the diskette button at the top of the view labeled Save the current expression. Both the expression and the test cases will be saved to a file. The default extension is .regx.
Save as button - To save the current expression with a different file name than what has previously been saved, click the button at the top of the view labeled Save the current expression as.
This section provides charts of regular expressions for each supported syntax (Brief, Perl, SlickEdit, UNIX, and Wildcards), including examples.
There are some differences between our implementation of these syntaxes and those elsewhere.
Brief regular expressions are defined in the following table.
|
Brief Regular Expression |
Definition |
|---|---|
|
% |
Matches beginning of line. |
|
< |
Matches beginning of line. |
|
$ |
Matches end of line. |
|
> |
Matches end of line. |
|
? |
Matches any character except newline. |
|
* |
Minimal match of zero or more of any character except newline. This is the same as ?@. |
|
X+ |
Minimal match of one or more occurrences of X. See Minimal versus Maximal Matching for more information. |
|
X\:* |
Maximal match of zero or more of any character except newline. This is the same as ?\:@. |
|
X\:@ |
Maximal match of zero or more occurrences of X. |
|
X\:+ |
Maximal match of one or more occurrences of X. |
|
X\: n1 |
Matches exactly n1 occurrences of X. Use {} to avoid ambiguous expressions. For example, a:9{}1 searches for nine instances of the letter "a" followed by a "1". |
|
X\: n1 , |
Maximal match of at least n1 occurrences of X. |
|
X\:, n2 |
Maximal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X\: n1 , n2 |
Maximal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
X\: n1 ? |
Match exactly n1 occurrences of X. |
|
X\: n1 ,? |
Minimal match of at least n1 occurrences of X. |
|
X\:, n2 ? |
Minimal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X\: n1 , n2 ? |
Minimal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
\(X\) |
Matches subexpression X but does not define a tagged expression. |
|
{X} |
Matches subexpression X and specifies a new tagged expression. See Using Tagged Expressions for more information. |
|
{@ d X} |
Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax {X} once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions. |
|
X|Y |
Matches X or Y. |
|
~{X} |
Search fails if expression X is matched. |
|
[ char-set ] |
Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. Backslash (\) can be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\0-\27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket. |
|
[~ char-set ] |
Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. The expression [~A-Z] matches all characters except uppercase letters. The expression [~] matches any character except newline. |
|
[ char-set1 - [ char-set2 ]] |
Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. For example, [a-z-[qw]] matches all English lowercase letters except "q" and "w". [\p{L}-[qw]] matches all Unicode lowercase letters except "q" and "w". |
|
[ char-set1 & [ char-set2 ]] |
Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127. |
|
\x{ hhhh } |
Matches up to 31-bit Unicode hexadecimal character specified by hhhh. |
|
\p{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\P{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\p{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions. |
|
\P{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions. |
|
\x hh |
Matches hexadecimal character hh where 0<=hh<=0xff. |
|
\d ddd |
Matches decimal character ddd where 0<=ddd<=255. |
|
\ d |
Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails. |
|
\c |
Specifies cursor position if match is found. If the expression xyz\c is found, the cursor is placed after z. |
|
\n |
Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file. Use \d10 if you want to match a 10 character. |
|
\r |
Matches carriage return. |
|
\t |
Matches tab character. |
|
\b |
Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. Note that this notation previously matched a backspace character. It can still be used to match a backspace character by using it in a character set (for example, [\b]). |
|
\B |
Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word. |
|
\Q and \E |
\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching. |
|
\f |
Matches form feed character. |
|
\od |
Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [~\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing. |
|
\om |
Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om?\@ matches the rest of the buffer. |
|
\ol |
Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like ?\@ will not match multiple lines. This is much safer and usually faster than using the \om option. |
|
\oi |
Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oc. |
|
\oc |
Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oi. |
|
\ char |
Declares character after slash to be literal. For example, \* represents the asterisk (*) character. |
|
\: char |
Matches predefined expression corresponding to char. The predefined expressions are:
Warning\:f and \:p Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed. Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename. |
The table below shows examples of Brief regular expressions.
|
Brief Regular Expression Example |
Description |
|---|---|
|
<defproc |
Matches lines that begin with the word defproc. |
|
<definit> |
Matches lines that only contain the word definit. |
|
<\*name |
Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *. |
|
[\t ] |
Matches tab and space characters. |
|
[\d9\d32] |
Matches tab and space characters. |
|
[\x9\x20] |
Matches tab and space characters. |
|
p?t |
Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat. |
|
s*t |
Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st. |
|
{for}|{while} |
Matches the strings for or while. |
|
^\:p |
Matches lines beginning with a file name. |
|
xy+z |
Matches x followed by one or more occurrences of y followed by z. |
|
\x0d\x0a\x01\x02 | |
|
\d13\d10\d1\d2 |
Matches a sequence of decimal binary characters. |
Perl regular expressions are defined in the following table.
|
Perl Regular Expression |
Definition |
|---|---|
|
^ |
Matches beginning of line. |
|
$ |
Matches end of line. |
|
. |
Matches any character except newline. |
|
X+ |
Maximal match of one or more occurrences of X. See Minimal versus Maximal Matching. |
|
X* |
Maximal match of zero or more occurrences of X. |
|
X? |
Maximal match of zero or one occurrences of X. |
|
X{ n1 } |
Match exactly n1 occurrences of X. |
|
X{ n1 ,} |
Maximal match of at least n1 occurrences of X. |
|
X{, n2 } |
Maximal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X{ n1 , n2 } |
Maximal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
X+? |
Minimal match of one or more occurrences of X. |
|
X*? |
Minimal match of zero or more occurrences of X. |
|
X?? |
Minimal match of zero or one occurrences of X. |
|
X{ n1 }? |
Matches exactly n1 occurrences of X. |
|
X{ n1 ,}? |
Minimal match of at least n1 occurrences of X. |
|
X{, n2 }? |
Minimal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X{ n1 , n2 }? |
Minimal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
(?!X) |
Search fails if expression X is matched. The expression ^(?!if) matches the beginning of all lines that do not start with if. |
|
(?=X) |
Assert, positive lookahead. Searches for subexpression X, but X is not returned as part of the match. For example, to match words ending in "ed" while excluding "ed" as part of the match, use \b[a-z]+(?=ed\b). See also (?!X). |
|
(?>X) |
Prohibit backtracking. This expression is advanced usage. It can be used to prevent the subexpression X from backtracking when using maximal (greedy) matching. |
|
(?#text) |
Comment. No text is matched in this expression; it is used for comment and documentation only. |
|
(X) |
Matches subexpression X and specifies a new tagged expression (see Using Tagged Expressions). No more tagged expressions are defined once an explicit tagged expression number is specified as shown below. |
|
(? d X) |
Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax (X) once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions. |
|
(?:X) |
Matches subexpression X but does not define a tagged expression. |
|
X|Y |
Matches X or Y. |
|
[ char-set ] |
Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. A backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\d0-\d27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket. The expression [^] matches a caret (^) character but this does not work for SlickEdit regular expressions. In both syntaxes, [\^] matches a caret (^) character. |
|
[^ char-set ] |
Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. |
|
[ char-set1 - [ char-set2 ]] |
Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. The expression [^A-Z] matches all characters except uppercase letters. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w. |
|
[ char-set1 & [ char-set2 ] |
Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127. |
|
\x{ hhhh } |
Matches up to 31-bit Unicode hexadecimal character specified by hhhh. |
|
\p{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\P{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\p{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions. |
|
\P{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions. |
|
\x hh |
Matches hexadecimal character hh where 0<=hh<=0xff. |
|
\#ddd |
Matches decimal character where 0<=ddd<=255. |
|
\ d |
Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails. |
|
\d |
Equivalent to [0-9]. Can also be used inside a character class. For example, [A-F\d] is equivalent to [A-F0-9]. |
|
\D |
Equivalent to [^0-9]. Can also be used inside a character class. |
|
\w |
Equivalent to [a-zA-Z0-9_]. Can also be used inside a character class. |
|
\W |
Equivalent to [^a-zA-Z0-9_]. Can also be used inside a character class. |
|
\s |
Equivalent to [ \t\n\r\f]. Can also be used inside a character class. |
|
\S |
Equivalent to [^ \t\n\r\f]. Can also be used inside a character class. |
|
\ooo |
Octal ASCII value. |
|
\cx |
Control character (ASCII values 0-31) '@' <=x<='_' |
|
\z |
Specifies cursor position if match is found. If the expression abc\z is found, the cursor is placed after the c. Note that in UNIX, this is the same as \c. However in Perl, \c is used only for control characters. |
|
\n |
Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character. |
|
\r |
Matches carriage return (ASCII 13). What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file. |
|
\t |
Matches tab character. |
|
\b |
Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. |
|
\B |
Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word. |
|
\Q and \E |
\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching. |
|
\f |
Matches form feed character. |
|
\od |
Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing. |
|
\om |
Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om.+ matches the rest of the buffer. |
|
\ol |
Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like .+ will not match multiple lines. This is much safer and usually faster than using the \om option. |
|
\oi |
Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. Note that this is the equivalent to the Perl syntax ?i. See also \oc. |
|
\oc |
Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. Note that this is equivalent to the Perl syntax ?-i. See also \oi. |
|
\ char |
Declares character after slash to be literal. For example, \* represents the star character. |
|
\: char |
Matches predefined expression corresponding to char. The pre-defined expressions are:
Warning\:f and \:p Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed. Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename. |
The precedence of operators, from highest to lowest, is as follows:
+, *, ?, {}, +?, *?, ??, {}? (These operators have the same precedence.)
concatenation
|
The table below shows examples of Perl regular expressions.
|
Perl Regular Expression Example |
Description |
|---|---|
|
^defproc |
Matches lines that begin with the word defproc. |
|
^definit$ |
Matches lines that only contain the word definit. |
|
^\*name |
Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *. |
|
[\t ] |
Matches tab and space characters. |
|
[\d9\d32] |
Matches tab and space characters. |
|
[\x9\x20] |
Matches tab and space characters. |
|
p.t |
Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat. |
|
s.*?t |
Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st. |
|
for|while |
Matches the strings for or while. |
|
^\:p |
Matches lines beginning with a file name. |
|
xy+z |
Matches x followed by one or more occurrences of y followed by z. |
|
[a-z-[qw]] |
Character set subtraction. Matches all English lowercase letters except q and w. |
|
[\p{isGreek}&[\p{L}]] |
Character set intersection. Matches all Unicode letters in the Greek block. |
|
\x{6587} |
Matches Unicode character with hexadecimal value 6587. Character set intersection. Matches all Unicode letters in the Greek block. |
|
[\p{L}-[qw]] |
Matches all Unicode letters except q and w. |
|
[\p{L}] |
Matches all Unicode letters. |
|
[\p{Lul}] |
Matches all Unicode uppercase and lowercase letters. |
|
[\P{L}] |
Matches all Unicode characters that are not letters. |
|
[\p{isGreek}] |
Matches all Unicode characters in the Greek block. |
|
\x0d\x0a\x01\x02 |
Matches a sequence of hex binary characters. |
|
\d13\d10\d1\d2 |
Matches a sequence of decimal binary characters. |
SlickEdit regular expressions are defined in the following table.
|
SlickEdit Regular Expression |
Definition |
|---|---|
|
^ |
Matches beginning of line. |
|
$ |
Matches end of line. |
|
? |
Matches any character except newline. |
|
X+ |
Minimal match of one or more occurrences of X. See Minimal versus Maximal Matching for more information. |
|
X# |
Maximal match of one or more occurrences of X. |
|
X* |
Minimal match of zero or more occurrences of X. |
|
X@ |
Maximal match of zero or more occurrences of X. |
|
X: n1 |
Matches exactly n1 occurrences of X. Use () to avoid ambiguous expressions. For example a:9()1 searches for nine instance of the letter a followed by a 1. |
|
X: n1 , |
Maximal match of at least n1 occurrences of X. |
|
X: n1 , n2 |
Maximal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
X:* n1 , |
Minimal match of at least n1 occurrences of X. |
|
X:* n1 , n2 |
Minimal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
~X |
Search fails if expression X is matched. The expression ^~(if) matches the beginning of all lines that do not start with if. |
|
(X) |
Matches subexpression X. |
|
{X} |
Matches subexpression X and specifies a new tagged expression. See Using Tagged Expressions for more information. |
|
{# d X} |
Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. |
|
X|Y |
Matches X or Y. |
|
[ char-set ] |
Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. Backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\0-\27] matches ASCII character codes 0..27. The expression [] matches no characters. In UNIX regular expressions, []] matches a right bracket. In both syntaxes, the expression [\]] matches a right bracket. The expression [\^] matches a caret (^) character in both syntaxes. |
|
[~ char-set ] |
Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. The expression [~A-Z] matches all characters except uppercase letters. |
|
[^ char-set ] |
Same as [~char-set] above. |
|
[ char-set1 - [ char-set2 ]] |
Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w. |
|
[ char-set1 & [ char-set2 ]] |
Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127. |
|
\x{ hhhh } |
Matches up to 31-bit Unicode hexadecimal character specified by hhhh. |
|
\p{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\P{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\p{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions. |
|
\P{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions. |
|
\x hh |
Matches hexadecimal character hh where 0<=hh<=0xff. |
|
\ ddd |
Matches decimal character ddd where 0<=ddd<=255. |
|
\g d |
Defines a back reference to tagged expression number d. For example, {abc}def\g0 matches the string abcdefabc. If the tagged expression has not been set, the search fails. |
|
\c |
Specifies cursor position if match is found. If the expression xyz\c is found, the cursor is placed after the z. |
|
\n |
Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character. |
|
\r |
Matches carriage return. |
|
\t |
Matches tab character. |
|
\b |
Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. Note that this notation previously matched a backspace character. It can still be used to match a backspace character by using it in a character set (for example, [\b]). |
|
\B |
Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word. |
|
\Q and \E |
\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching. |
|
\f |
Matches form feed character. |
|
\od |
Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing. |
|
\om |
Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om?# matches the rest of the buffer. NOTE: Test the regular expression on a very small file before using it on a large file. This option may cause the editor to use a lot of memory. |
|
\ol |
Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like ?# will not match multiple lines. This is much safer and usually faster than using the \om option. |
|
\oi |
Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside all ( ) and { } groups, after which case matching is restored to the previous case match setting. See also \oc. |
|
\oc |
Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside all ( ) and { } groups, after which case matching is restored to the previous case match setting. See also \oi. |
|
\ char |
Declares character after slash to be literal. For example, \: represents the colon character. |
|
: char |
Matches predefined expression corresponding to char. The predefined expressions are:
Warning:f and :p Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed. Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename. |
The precedence of operators, from highest to lowest, is as follows:
+, #, *, @, :, :* (These operators have the same precedence.)
concatenation
|
The table below shows examples of SlickEdit regular expressions.
|
SlickEdit Regular Expression Example |
Description |
|---|---|
|
^defproc |
Matches lines that begin with the word defproc. |
|
^definit$ |
Matches lines that only contain the word definit. |
|
^\:name |
Matches lines that begin with the string :name. Notice that the backslash must prefix the colon character (:). |
|
[\t ] |
Matches tab and space characters. |
|
[\9\32] |
Matches tab and space characters. |
|
[\x9\x20] |
Matches tab and space characters. |
|
p?t |
Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat. |
|
s?*t |
Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st. |
|
for|while |
Matches the strings for or while. |
|
^:p |
Matches lines beginning with a file name. |
|
xy+z |
Matches x followed by one or more occurrences of y followed by z. |
|
\x0d\x0a\x01\x02 |
Matches a sequence of hex binary characters. |
|
\13\10\1\2 |
Matches a sequence of decimal binary characters. |
UNIX regular expressions are defined in the following table.
|
UNIX Regular Expression |
Definition |
|---|---|
|
^ |
Matches beginning of line. |
|
$ |
Matches end of line. |
|
. |
Matches any character except newline. |
|
X+ |
Maximal match of one or more occurrences of X. See Minimal versus Maximal Matching. |
|
X* |
Maximal match of zero or more occurrences of X. |
|
X? |
Maximal match of zero or one occurrences of X. |
|
X{ n1 } |
Match exactly n1 occurrences of X. |
|
X{ n1 ,} |
Maximal match of at least n1 occurrences of X. |
|
X{, n2 } |
Maximal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X{ n1 , n2 } |
Maximal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
X+? |
Minimal match of one or more occurrences of X. |
|
X*? |
Minimal match of zero or more occurrences of X. |
|
X?? |
Minimal match of zero or one occurrences of X. |
|
X{ n1 }? |
Matches exactly n1 occurrences of X. |
|
X{ n1 ,}? |
Minimal match of at least n1 occurrences of X. |
|
X{, n2 }? |
Minimal match of at least zero occurrences but not more than n2 occurrences of X. |
|
X{ n1 , n2 }? |
Minimal match of at least n1 occurrences but not more than n2 occurrences of X. |
|
(?!X) |
Search fails if expression X is matched. The expression ^(?!if) matches the beginning of all lines that do not start with if. |
|
(?=X) |
Assert, positive lookahead. Searches for subexpression X, but X is not returned as part of the match. For example, to match words ending in "ed" while excluding "ed" as part of the match, use \b[a-z]+(?=ed\b). See also (?!X). |
|
(?>X) |
Prohibit backtracking. This expression is advanced usage. It can be used to prevent the subexpression X from backtracking when using maximal (greedy) matching. |
|
(?#text) |
Comment. No text is matched in this expression; it is used for comment and documentation only. |
|
(X) |
Matches subexpression X and specifies a new tagged expression (see Using Tagged Expressions). No more tagged expressions are defined once an explicit tagged expression number is specified as shown below. |
|
(? d X) |
Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax (X) once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions. |
|
(?:X) |
Matches subexpression X but does not define a tagged expression. |
|
X|Y |
Matches X or Y. |
|
[ char-set ] |
Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. A backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\d0-\d27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket. The expression [^] matches a caret (^) character but this does not work for SlickEdit regular expressions. In both syntaxes, [\^] matches a caret (^) character. |
|
[^ char-set ] |
Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. |
|
[ char-set1 - [ char-set2 ]] |
Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. The expression [^A-Z] matches all characters except uppercase letters. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w. |
|
[ char-set1 & [ char-set2 ] |
Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127. |
|
\x{ hhhh } |
Matches up to 31-bit Unicode hexadecimal character specified by hhhh. |
|
\p{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\P{ UnicodeCategorySpec ] |
(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions. |
|
\p{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions. |
|
\P{ UnicodeIsBlockSpec ] |
(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions. |
|
\x hh |
Matches hexadecimal character hh where 0<=hh<=0xff. |
|
\d ddd |
Matches decimal character ddd where 0<=ddd<=255. |
|
\ d |
Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails. |
|
\c |
Specifies cursor position if match is found. If the expression xyz\c is found the cursor is placed after the z. |
|
\n |
Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character. |
|
\r |
Matches carriage return (ASCII 13). What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file. |
|
\t |
Matches tab character. |
|
\b |
Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. |
|
\B |
Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word. |
|
\Q and \E |
\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching. |
|
\f |
Matches form feed character. |
|
\od |
Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing. |
|
\om |
Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om.+ matches the rest of the buffer. |
|
\ol |
Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like .+ will not match multiple lines. This is much safer and usually faster than using the \om option. |
|
\oi |
Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oc. |
|
\oc |
Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oi. |
|
\ char |
Declares character after slash to be literal. For example, \* represents the star character. |
|
\: char |
Matches predefined expression corresponding to char. The pre-defined expressions are:
Warning\:f and \:p Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed. Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename. |
The precedence of operators, from highest to lowest, is as follows:
+, *, ?, {}, +?, *?, ??, {}? (These operators have the same precedence.)
concatenation
|
The table below shows examples of UNIX regular expressions.
|
UNIX Regular Expression Example |
Description |
|---|---|
|
^defproc |
Matches lines that begin with the word defproc. |
|
^definit$ |
Matches lines that only contain the word definit. |
|
^\*name |
Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *. |
|
[\t ] |
Matches tab and space characters. |
|
[\d9\d32] |
Matches tab and space characters. |
|
[\x9\x20] |
Matches tab and space characters. |
|
p.t |
Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat. |
|
s.*?t |
Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st. |
|
for|while |
Matches the strings for or while. |
|
^\:p |
Matches lines beginning with a file name. |
|
xy+z |
Matches x followed by one or more occurrences of y followed by z. |
|
[a-z-[qw]] |
Character set subtraction. Matches all English lowercase letters except q and w. |
|
[\p{isGreek}&[\p{L}]] |
Character set intersection. Matches all Unicode letters in the Greek block. |
|
\x{6587} |
Matches Unicode character with hexadecimal value 6587. Character set intersection. Matches all Unicode letters in the Greek block. |
|
[\p{L}-[qw]] |
Matches all Unicode letters except q and w. |
|
[\p{L}] |
Matches all Unicode letters. |
|
[\p{Lul}] |
Matches all Unicode uppercase and lowercase letters. |
|
[\P{L}] |
Matches all Unicode characters that are not letters. |
|
[\p{isGreek}] |
Matches all Unicode characters in the Greek block. |
|
\x0d\x0a\x01\x02 |
Matches a sequence of hex binary characters. |
|
\d13\d10\d1\d2 |
Matches a sequence of decimal binary characters. |
SlickEdit® Core supports *, ?, and # wildcards:
The asterisk (*) matches zero or more characters. For example, search for a*b to find any string that contains a lowercase letter "a" followed by a lowercase letter "b" allowing for text in between.
The question mark (?) matches any single character. Use multiple question marks in succession to represent that number of characters. For example, search for a???b to find any string that contains a lowercase letter "a" followed by any three characters, followed by a lowercase letter "b".
The pound sign (#) matches any single digit, 0-9. Use multiple pound signs in succession to represent that number of digits. For example, use ##:## to search for four-digit time-of-day values.
The Unicode consortium standard regular expression categories are supported. The syntax for specifying categories is:
\p{MainCategoryLetter Subcategories}The above syntax matches the categories specified. The following syntax matches all characters not in the categories specified:
\P{MainCategoryLetter Subcategories}The \p and \P notations can only be used inside a character set specification. MainCategoryLetter can be L, M, N, P, S, Z, or C. The valid Subcategories depend on the MainCategoryLetter specified. If no Subcategories are specified, all are assumed. For example:
[\p{L}] matches all Unicode letters.
[\p{Lul}] matches all uppercase and lowercase letters.
[\P{L}] matches all characters that are not letters.
The following table lists the valid subcategories for a specific main category. These character tables were generated using the file UnicodeData-3.1.0.txt found on the Unicode Consortium Web site
(http://unicode.org).
|
Subcategory |
Description |
|---|---|
|
Lu |
Letter, Uppercase |
|
Ll |
Letter, Lowercase |
|
Lt |
Letter, Titlecase |
|
Lo |
Letter, Other |
|
Mn |
Mark, Non-Spacing |
|
Mc |
Mark, Spacing Combining |
|
Me |
Mark, Enclosing |
|
Nd |
Number, Decimal Digit |
|
Nl |
Number, Letter |
|
No |
Number, Other |
|
Pc |
Punctuation, Connector |
|
Pd |
Punctuation, Dash |
|
Ps |
Punctuation, Open |
|
Pe |
Punctuation, Close |
|
Pi |
Punctuation, Initial quote (may behave like Ps or Pe depending on usage) |
|
Pf |
Punctuation, Final quote (may behave like Ps or Pe depending on usage) |
|
Po |
Punctuation, Other |
|
Sm |
Symbol, Math |
|
Sc |
Symbol, Currency |
|
Sk |
Symbol, Modifier |
|
So |
Symbol, Other |
|
Zs |
Separator, Space |
|
Zl |
Separator, Line |
|
Zp |
Separator, Paragraph |
|
Cc |
Other, Control |
|
Cf |
Other, Format |
|
Cs |
Other, Surrogate |
|
Co |
Other, Private Use |
|
Cn |
Other, Not Assigned (no characters in the file have this property) |
The Unicode consortium standard regular expression block categories are supported. The syntax for specifying a character block is:
\p{Is BlockName}The above syntax matches the characters in the block specified. The following syntax matches all characters not in the block specified:
\P{Is BlockName }The \p and \P notations may only be used inside a character set specification. For example, [\p{isBasicLatin}] matches all characters in the Greek block. [\P{isBasicLatin}] matches all characters that are not in the Greek block.
The following table lists the non-standard valid character block names. These character tables were generated from XML standards found at the World Wide Web Consortium Web site).
|
Block Name |
Description |
|---|---|
|
XMLNameStartChar |
All characters that are valid for the start of an XML tag name. |
|
XMLNameChar |
All characters that are valid in an XML tag name. |
The following table lists the valid character block names. These character tables were generated using the blocks.txt file found on the Unicode Consortium Web site (http://unicode.org).
|
Range |
Block Name |
|---|---|
|
0000..007F |
BasicLatin |
|
0080..00FF |
Latin-1Supplement |
|
0100..017F |
LatinExtended-A |
|
0180..024F |
LatinExtended-B |
|
0250..02AF |
IPAExtensions |
|
02B0..02FF |
SpacingModifierLetters |
|
0300..036F |
CombiningDiacriticalMarks |
|
0370..03FF |
Greek |
|
0400..04FF |
Cyrillic |
|
0530..058F |
Armenian |
|
0590..05FF |
Hebrew |
|
0600..06FF |
Arabic |
|
0700..074F |
Syriac |
|
0780..07BF |
Thaana |
|
0900..097F |
Devanagari |
|
0980..09FF |
Bengali |
|
0A00..0A7F |
Gurmukhi |
|
0A80..0AFF |
Gujarati |
|
0B00..0B7F |
Oriya |
|
0B80..0BFF |
Tamil |
|
0C00..0C7F |
Telugu |
|
0C80..0CFF |
Kannada |
|
0D00..0D7F |
Malayalam |
|
0D80..0DFF |
Sinhala |
|
0E00..0E7F |
Thai |
|
0E80..0EFF |
Lao |
|
0F00..0FFF |
Tibetan |
|
1000..109F |
Myanmar |
|
10A0..10FF |
Georgian |
|
1100..11FF |
HangulJamo |
|
1200..137F |
Ethiopic |
|
13A0..13FF |
Cherokee |
|
1400..167F |
UnifiedCanadianAboriginalSyllabics |
|
1680..169F |
Ogham |
|
16A0..16FF |
Runic |
|
1780..17FF |
Khmer |
|
1800..18AF |
Mongolian |
|
1E00..1EFF |
LatinExtendedAdditional |
|
1F00..1FFF |
GreekExtended |
|
2000..206F |
GeneralPunctuation |
|
2070..209F |
SuperscriptsandSubscripts |
|
20A0..20CF |
CurrencySymbols |
|
20D0..20FF |
CombiningMarksforSymbols |
|
2100..214F |
LetterlikeSymbols |
|
2150..218F |
NumberForms |
|
2190..21FF |
Arrows |
|
2200..22FF |
MathematicalOperators |
|
2300..23FF |
MiscellaneousTechnical |
|
2400..243F |
ControlPictures |
|
2440..245F |
OpticalCharacterRecognition |
|
2460..24FF |
EnclosedAlphanumerics |
|
2500..257F |
BoxDrawing |
|
2580..259F |
BlockElements |
|
25A0..25FF |
GeometricShapes |
|
2600..26FF |
MiscellaneousSymbols |
|
2700..27BF |
Dingbats |
|
2800..28FF |
BraillePatterns |
|
2E80..2EFF |
CJKRadicalsSupplement |
|
2F00..2FDF |
KangxiRadicals |
|
2FF0..2FFF |
IdeographicDescriptionCharacters |
|
3000..303F |
CJKSymbolsandPunctuation |
|
3040..309F |
Hiragana |
|
30A0..30FF |
Katakana |
|
3100..312F |
Bopomofo |
|
3130..318F |
HangulCompatibilityJamo |
|
3190..319F |
Kanbun |
|
31A0..31BF |
BopomofoExtended |
|
3200..32FF |
EnclosedCJKLettersandMonths |
|
3300..33FF |
CJKCompatibility |
|
3400..4DB5 |
CJKUnifiedIdeographsExtensionA |
|
4E00..9FFF |
CJKUnifiedIdeographs |
|
A000..A48F |
YiSyllables |
|
A490..A4CF |
YiRadicals |
|
AC00..D7A3 |
HangulSyllables |
|
D800..DB7F |
HighSurrogates |
|
DB80..DBFF |
HighPrivateUseSurrogates |
|
DC00..DFFF |
LowSurrogates |
|
E000..F8FF |
PrivateUse |
|
F900..FAFF |
CJKCompatibilityIdeographs |
|
FB00..FB4F |
AlphabeticPresentationForms |
|
FB50..FDFF |
ArabicPresentationForms-A |
|
FE20..FE2F |
CombiningHalfMarks |
|
FE30..FE4F |
CJKCompatibilityForms |
|
FE50..FE6F |
SmallFormVariants |
|
FE70..FEFE |
ArabicPresentationForms-B |
|
FEFF..FEFF |
Specials |
|
FF00..FFEF |
HalfwidthandFullwidthForms |
|
FFF0..FFFD |
Specials |
|
10300..1032F |
OldItalic |
|
10330..1034F |
Gothic |
|
10400..1044F |
Deseret |
|
1D000..1D0FF |
ByzantineMusicalSymbols |
|
1D100..1D1FF |
MusicalSymbols |
|
1D400..1D7FF |
MathematicalAlphanumericSymbols |
|
20000..2A6D6 |
CJKUnifiedIdeographsExtensionB |
|
2F800..2FA1F |
CJKCompatibilityIdeographsSupplement |
|
E0000..E007F |
Tags |