Regular Expressions

Overview

Regular expressions are patterns of text used to match and manipulate strings in your code. These patterns are expressed with combinations of characters defined by the regular expression syntax being used. A regular expression is sometimes referred to as a "regex".

Use regular expressions in your search and replace operations when you find normal search/replace too limiting. For example, with regular expressions, you can:

  • Find quoted strings.

  • Find blank lines.

  • Find words starting at the beginning of lines.

  • Find two words separated by any number of spaces or other text.

SlickEdit® Core supports five types of regular expression syntax:

SlickEdit Core also provides a Regex Evaluator that you can use to interactively create, save, and re-use tests of regular expressions. See The Regex Evaluator for more information.

See Regular Expression Syntax for charts of the expressions in syntax. Unicode regular expression categories and character blocks are also supported. See Unicode Categories and Character Blocks for more information.

Note

  • This documentation is not meant to be an exhaustive resource on regular expressions. Rather, we will present basic information, syntax charts, and examples. For novice users, there are many books and Web sites that go into more detail about this topic.

  • While regular expressions in SlickEdit Core primarily match the syntax for that language, there are some differences between our implementation and those elsewhere.

Using Regular Expressions in SlickEdit® Core

Specifying the Syntax to Use

All search and replace commands, the Find and Replace view, and incremental search support regular expressions. For search and replace commands and the view, you can specify the regular expression syntax to use through specific options. A global option is available to specify the default syntax to use when you invoke these features or when you use incremental search.

For example:

  • Search and replace commands - When using the search commands / (slash) and find, or the replace commands c and replace, you can use the following options to specify regular expression syntax:

    • Use B to interpret the search string as Brief regular expression.

    • Use the L option to interpret the string as a Perl regular expression.

    • Use R to interpret the string as a SlickEdit regular expression.

    • Use the U option to interpret the string as a UNIX regular expression.

  • Find and Replace view - When using the view, select Use in the Search options box, and then pick the syntax to use from the drop-down list.

  • Incremental search - When using incremental search, press Ctrl+T to toggle regular expression searching on and off. The syntax that will be used is based on the global syntax setting.

To set the global option, from the main menu, click Window → SlickEdit Preferences, expand Editing and select Search. Set the Regular expression option to True and select the syntax you want to use from the Expression type drop-down list.

Minimal versus Maximal Matching

If you are using tagged expressions or regular expressions to perform a search and replace, it is important to understand the difference between the minimal and maximal operators.

Take, for example, a line of text which contains a DOS file name: \path1\path2\path3\name.ext.

Based on the syntax, the following regular expressions match the string \path1\:

Syntax

Expression

Brief

<\\*\\

Perl and UNIX

^\\.*?\\

SlickEdit

^\\?*\\

The following regular expressions, which use the maximal operator, match the string \path\path2\path3\:

Syntax

Expression

Brief

<\\\:*\\

Perl and UNIX

^\\.*\\

SlickEdit

^\\?@\\

As a rule of thumb, the following minimal matching operators are generally used after a less-specific regular expression such as ? in Brief/SlickEdit or .in UNIX:

Syntax

Operators

Brief

@ and +

Perl and UNIX

*? and +?

SlickEdit

* and +

Use the maximal matching operators after a regular expression which matches something more specific. For example, to search for a string of digits and prefix each matched string with the character $, specify the following expressions:

Syntax

Expression

Brief

Search for: {[0-9]\:+}

Replace with:$\0

Perl and UNIX

Search for:([0-9]+)

Replace with:$\1

SlickEdit

Search for: {[0-9]#}

Replace with: $#0

If the minimal matching operator (+ in Brief/SlickEdit syntax, +? in Perl/UNIX) was used in the search string instead of the maximal matching operator (\:@ in Brief, + in Perl/UNIX, # in SlickEdit Core), the above search and replace would prefix each digit in the entire file with a $ character.

Using Tagged Expressions

When you use regular expressions to search for a string, you will often want the replace string to depend on what was found. Use tagged expressions to insert parts of what is found into the replace string. Tagged expressions are denoted based on the syntax:

  • Brief syntax - Use { } (curly braces) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) followed by a tagged expression number 0-9. Count the { (left braces)in the search string to determine a tagged expression number. The first tagged expression is \0.

  • UNIX syntax - Use ( ) (parentheses) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) followed by a tag group number 1-9. Count the ( (left parenthesis) in the search string to determine a tagged expression number. The first tagged expression is \1 and the last is \0.

  • Perl syntax - Use ( ) (parentheses) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a \ (backslash) or $ (dollar sign) followed by a tag group number 1-9. Count the ( (left parenthesis) in the search string to determine a tagged expression number. The first tagged expression is \1 and the last is \0.

  • SlickEdit® Core syntax - Use { } (curly braces) to denote a tagged expression in the search string. The replace string specifies tagged expressions with a # (pound sign) followed by a tagged expression number 0-9. Count the { (left braces) in the search string to determine a tagged expression number. The first tagged expression is #0.

Examples of Tagged Expressions

Example 1: Replace Occurrences

The expressions in the table below replace occurrences of "if" and "while" with "xify" and "xwhiley." Unmatched groups are null. Note that the \1 in Brief syntax, \2 in Perl/UNIX, and #1 in SlickEdit syntax are replaced with null.

Syntax

Expression

Brief

Search for: {{if}|while}}

Replace with: x\0y\1

Perl/UNIX

Search for: (if|while)

Replace with: x\1y\2

SlickEdit

Search for: {if|while}

Replace with:x#0y#1

Example 2: Reverse Text on Lines Containing a Comma

The expressions in the table below reverse text on lines containing a comma. Lines with "abc,def" will be changed to "def,abc". Notice that the Perl/UNIX regular expression search string uses a *? minimal matching operator, so the comma actually matches the first comma in the line and not the last.

Syntax

Expression

Brief

Search for:^{*},{*}$

Replace with: \1,\0

Perl/UNIX

Search for: ^(.*?),(.*)$

Replace with: \2,\1

SlickEdit

Search for:^{?*},{?*}$

Replace with:#1,#0

Replacing with Regular Expressions

When using regular expressions, some characters have a different meaning when used in the replace string, depending on the syntax:

  • Brief - The backslash character (\) in the replace string has the same meaning as in the search string except that \c and \:char are not supported.

  • UNIX - A backslash in the replace string has the same meaning as in the search string except that \c and \:char are not supported.

  • Perl - A backslash in the replace string has the same meaning as in the search string except that \c and \:char are not supported. A dollar sign ($) must be escaped (\$) when replacing a literal $.

  • SlickEdit® Core - The pound sign character (#) and backslash (\) have special meaning in the replace string. A backslash in the replace string has the same meaning as in the search string except that \c, :char, and \gd are not supported.

See the Regular Expression Syntax tables for a list of options for these characters. See Using Tagged Expressions for information on specifying tagged expressions in the replace string.

Case Modification in Replace

When used in a replace operation, the expressions in the following table can be used to modify the character casing of matched expressions. These work in Brief, Perl, SlickEdit, and UNIX syntaxes.

Expression

Description

\l

Convert next character lowercase.

\u

Convert next character uppercase.

\L

Convert all characters lowercase until \E.

\U

Convert all characters uppercase until \E.

\Q

Replace all characters literally until \E.

\E

End all case modification or \Q.

Examples of Replacing with Regular Expressions

The table below contains some examples of replace operations using regular expressions.

Operation

Expression

Search for occurrences of the string "hat" that occur at the end of a line and replace it with "cat".

In all syntaxes:

Search for: hat$

Replace with: cat

Delete blank lines.

Brief:

Search for: <\n

Replace with: (leave blank)

Perl/SlickEdit/UNIX:

Search for: ^\n

Replace with: (leave blank)

Replace occurrences of two consecutive blank lines with one.

Brief:

Search for: <\n\n

Replace with: \n

Perl/SlickEdit/UNIX:

Search for: ^\n\n

Replace with:\n

Search for lines containing "a" and replace the "a" with a formfeed character.

Brief:

Search for: <a+$

Replace with:\d12

Perl/UNIX:

Search for:^a+$

Replace with:\d12

SlickEdit:

Search for:^a+$

Replace with:\12

Select occurrences of "Title:" at the beginning of a line and capitalize the text following "Title:".

Brief:

Search for: <Title: {\:*}

Replace with: Title: \U\0

Perl/UNIX:

Search for: ^Title: (.*)

Replace with: Title: \U\1

SlickEdit:

Search for: ^Title\: {?@}

Replace with: Title: \U#0

The Regex Evaluator

Regular expressions are used to express text patterns for searching. The Regex Evaluator provides the capability to interactively create, save, and re-use tests of regular expressions.

To access the Regex Evaluator, click Tools → Regex Evaluator (or use the activate_regex_evaluator command). Like other views in SlickEdit® Core, this view is dockable. Docking options can be accessed by right-clicking on the view's title bar.

Type some samples of the text you are trying to match in the top portion of the view labeled Test Cases. Enter your regular expression pattern in the bottom field. The Regex Evaluator will highlight matched portions of your sample text and identify groups.

Entering Test Cases

Type your test cases in the Test Cases text box. These test cases will be evaluated as you type your regular expression in the bottom field. A wavy underline will indicate the ranges of text that match the entire expression. Matches are also marked with a yellow arrow that appears in the gutter to the left of the test case. You can hover your mouse on this arrow to see a tool tip which displays the matched expression details. When groups (tagged expressions) are used in your regular expression pattern, the groups will be boxed and highlighted in yellow in the Test Cases section.

Entering a Regular Expression

Enter the regular expression to test in the text field. Use the radio buttons to select the expression syntax that you wish to use: UNIX, SlickEdit® Core, Brief, or Perl. Click the arrow to the right of the regular expression field to pick from a menu of common syntax and operators.

Regex Evaluator Options

The following options and buttons are available on the Regex Evaluator view:

  • Multiline mode - If Multiline mode is selected, rather than searching through the test cases line-by-line, regular expressions will be searched on all lines at once. This is useful for test cases that wrap to the next line. This works just as if you had entered \om on the SlickEdit® Core command line.

  • Case sensitive - If Case sensitive is selected, the regular expression search will be case sensitive. This option is on by default.

  • New expression button - To clear the view of all entries in order to start a new evaluation, click the button at the top of the view labeled New expression.

  • Open a saved expression button - To open an expression that you have already saved, click the folder button at the top of the view labeled Open a saved expression.

  • Save the current expression button - To save the current expression, click the diskette button at the top of the view labeled Save the current expression. Both the expression and the test cases will be saved to a file. The default extension is .regx.

  • Save as button - To save the current expression with a different file name than what has previously been saved, click the button at the top of the view labeled Save the current expression as.

Regular Expression Syntax

This section provides charts of regular expressions for each supported syntax (Brief, Perl, SlickEdit, UNIX, and Wildcards), including examples.

Note

There are some differences between our implementation of these syntaxes and those elsewhere.

Brief Regular Expressions

Brief regular expressions are defined in the following table.

Brief Regular Expression

Definition

%

Matches beginning of line.

<

Matches beginning of line.

$

Matches end of line.

>

Matches end of line.

?

Matches any character except newline.

*

Minimal match of zero or more of any character except newline. This is the same as ?@.

X+

Minimal match of one or more occurrences of X. See Minimal versus Maximal Matching for more information.

X\:*

Maximal match of zero or more of any character except newline. This is the same as ?\:@.

X\:@

Maximal match of zero or more occurrences of X.

X\:+

Maximal match of one or more occurrences of X.

X\: n1

Matches exactly n1 occurrences of X. Use {} to avoid ambiguous expressions. For example, a:9{}1 searches for nine instances of the letter "a" followed by a "1".

X\: n1 ,

Maximal match of at least n1 occurrences of X.

X\:, n2

Maximal match of at least zero occurrences but not more than n2 occurrences of X.

X\: n1 , n2

Maximal match of at least n1 occurrences but not more than n2 occurrences of X.

X\: n1 ?

Match exactly n1 occurrences of X.

X\: n1 ,?

Minimal match of at least n1 occurrences of X.

X\:, n2 ?

Minimal match of at least zero occurrences but not more than n2 occurrences of X.

X\: n1 , n2 ?

Minimal match of at least n1 occurrences but not more than n2 occurrences of X.

\(X\)

Matches subexpression X but does not define a tagged expression.

{X}

Matches subexpression X and specifies a new tagged expression. See Using Tagged Expressions for more information.

{@ d X}

Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax {X} once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions.

X|Y

Matches X or Y.

~{X}

Search fails if expression X is matched.

[ char-set ]

Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. Backslash (\) can be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\0-\27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket.

[~ char-set ]

Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. The expression [~A-Z] matches all characters except uppercase letters. The expression [~] matches any character except newline.

[ char-set1 - [ char-set2 ]]

Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. For example, [a-z-[qw]] matches all English lowercase letters except "q" and "w". [\p{L}-[qw]] matches all Unicode lowercase letters except "q" and "w".

[ char-set1 & [ char-set2 ]]

Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127.

\x{ hhhh }

Matches up to 31-bit Unicode hexadecimal character specified by hhhh.

\p{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions.

\P{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions.

\p{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions.

\P{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions.

\x hh

Matches hexadecimal character hh where 0<=hh<=0xff.

\d ddd

Matches decimal character ddd where 0<=ddd<=255.

\ d

Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails.

\c

Specifies cursor position if match is found. If the expression xyz\c is found, the cursor is placed after z.

\n

Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file. Use \d10 if you want to match a 10 character.

\r

Matches carriage return.

\t

Matches tab character.

\b

Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. Note that this notation previously matched a backspace character. It can still be used to match a backspace character by using it in a character set (for example, [\b]).

\B

Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word.

\Q and \E

\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching.

\f

Matches form feed character.

\od

Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [~\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing.

\om

Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om?\@ matches the rest of the buffer.

\ol

Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like ?\@ will not match multiple lines. This is much safer and usually faster than using the \om option.

\oi

Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oc.

\oc

Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oi.

\ char

Declares character after slash to be literal. For example, \* represents the asterisk (*) character.

\: char

Matches predefined expression corresponding to char. The predefined expressions are:

  • \:a [A-Za-z0-9 - Matches an alphanumeric character.

  • \:b\([ \t]#\) - Matches blanks.

  • \:c [A-Za-z] - Matches an alphabetic character.

  • \:d [0-9] - Matches a digit.

  • \:f \([~\[\]\:\\/<>|=+;, \t"']#\) - Matches a file name part.

  • \:f \([~/ \t"']#\) - UNIX: Matches a file name part.

  • \:h\([0-9A-Fa-f]#\) - Matches a hex number.

  • \:i\([0-9]#\) - Matches an integer.

  • \:n\([0-9]#\(.[0-9]#|\)|.[0-9]#\)\([Ee]\(\+|-|\)[0-9]#|\)\) - Matches a floating number.

  • \:p\(\([A-Za-z]\:|\)\(\\|/|\)\(:f\(\\|/\)\)@:f\) - Windows: Matches a path.

  • \:p\(\(/|\)\(:f\(/\)\)@:f\) - UNIX: Matches a path.

  • \:q\(\"[~\"]@\"|'[~']@'\) - Matches a quoted string.

  • \:v\ ([A-Za-z_$][A-Za-z0-9_$]@\) - Matches a C variable.

  • \:w\ ([A-Za-z]#\) - Matches a word.

Warning

\:f and \:p

Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed.

Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename.

Brief Regular Expression Examples

The table below shows examples of Brief regular expressions.

Brief Regular Expression Example

Description

<defproc

Matches lines that begin with the word defproc.

<definit>

Matches lines that only contain the word definit.

<\*name

Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *.

[\t ]

Matches tab and space characters.

[\d9\d32]

Matches tab and space characters.

[\x9\x20]

Matches tab and space characters.

p?t

Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat.

s*t

Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st.

{for}|{while}

Matches the strings for or while.

^\:p

Matches lines beginning with a file name.

xy+z

Matches x followed by one or more occurrences of y followed by z.

\x0d\x0a\x01\x02

Matches a sequence of hex binary characters.

\d13\d10\d1\d2

Matches a sequence of decimal binary characters.

Perl Regular Expressions

Perl regular expressions are defined in the following table.

Perl Regular Expression

Definition

^

Matches beginning of line.

$

Matches end of line.

.

Matches any character except newline.

X+

Maximal match of one or more occurrences of X. See Minimal versus Maximal Matching.

X*

Maximal match of zero or more occurrences of X.

X?

Maximal match of zero or one occurrences of X.

X{ n1 }

Match exactly n1 occurrences of X.

X{ n1 ,}

Maximal match of at least n1 occurrences of X.

X{, n2 }

Maximal match of at least zero occurrences but not more than n2 occurrences of X.

X{ n1 , n2 }

Maximal match of at least n1 occurrences but not more than n2 occurrences of X.

X+?

Minimal match of one or more occurrences of X.

X*?

Minimal match of zero or more occurrences of X.

X??

Minimal match of zero or one occurrences of X.

X{ n1 }?

Matches exactly n1 occurrences of X.

X{ n1 ,}?

Minimal match of at least n1 occurrences of X.

X{, n2 }?

Minimal match of at least zero occurrences but not more than n2 occurrences of X.

X{ n1 , n2 }?

Minimal match of at least n1 occurrences but not more than n2 occurrences of X.

(?!X)

Search fails if expression X is matched. The expression ^(?!if) matches the beginning of all lines that do not start with if.

(?=X)

Assert, positive lookahead. Searches for subexpression X, but X is not returned as part of the match. For example, to match words ending in "ed" while excluding "ed" as part of the match, use \b[a-z]+(?=ed\b). See also (?!X).

(?>X)

Prohibit backtracking. This expression is advanced usage. It can be used to prevent the subexpression X from backtracking when using maximal (greedy) matching.

(?#text)

Comment. No text is matched in this expression; it is used for comment and documentation only.

(X)

Matches subexpression X and specifies a new tagged expression (see Using Tagged Expressions). No more tagged expressions are defined once an explicit tagged expression number is specified as shown below.

(? d X)

Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax (X) once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions.

(?:X)

Matches subexpression X but does not define a tagged expression.

X|Y

Matches X or Y.

[ char-set ]

Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. A backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\d0-\d27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket. The expression [^] matches a caret (^) character but this does not work for SlickEdit regular expressions. In both syntaxes, [\^] matches a caret (^) character.

[^ char-set ]

Matches any character not specified by char-set. A dash (-) character may be used to specify ranges.

[ char-set1 - [ char-set2 ]]

Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. The expression [^A-Z] matches all characters except uppercase letters. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w.

[ char-set1 & [ char-set2 ]

Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127.

\x{ hhhh }

Matches up to 31-bit Unicode hexadecimal character specified by hhhh.

\p{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions.

\P{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions.

\p{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions.

\P{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions.

\x hh

Matches hexadecimal character hh where 0<=hh<=0xff.

\#ddd

Matches decimal character where 0<=ddd<=255.

\ d

Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails.

\d

Equivalent to [0-9]. Can also be used inside a character class. For example, [A-F\d] is equivalent to [A-F0-9].

\D

Equivalent to [^0-9]. Can also be used inside a character class.

\w

Equivalent to [a-zA-Z0-9_]. Can also be used inside a character class.

\W

Equivalent to [^a-zA-Z0-9_]. Can also be used inside a character class.

\s

Equivalent to [ \t\n\r\f]. Can also be used inside a character class.

\S

Equivalent to [^ \t\n\r\f]. Can also be used inside a character class.

\ooo

Octal ASCII value.

\cx

Control character (ASCII values 0-31) '@' <=x<='_'

\z

Specifies cursor position if match is found. If the expression abc\z is found, the cursor is placed after the c. Note that in UNIX, this is the same as \c. However in Perl, \c is used only for control characters.

\n

Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character.

\r

Matches carriage return (ASCII 13). What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file.

\t

Matches tab character.

\b

Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word.

\B

Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word.

\Q and \E

\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching.

\f

Matches form feed character.

\od

Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing.

\om

Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om.+ matches the rest of the buffer.

\ol

Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like .+ will not match multiple lines. This is much safer and usually faster than using the \om option.

\oi

Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. Note that this is the equivalent to the Perl syntax ?i. See also \oc.

\oc

Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. Note that this is equivalent to the Perl syntax ?-i. See also \oi.

\ char

Declares character after slash to be literal. For example, \* represents the star character.

\: char

Matches predefined expression corresponding to char. The pre-defined expressions are:

  • \:a [A-Za-z0-9] - Matches an alphanumeric character.

  • \:c [A-Za-z] - Matches an alphabetic character.

  • \:b (?:[ \t]+) - Matches blanks.

  • \:d [0-9] - Matches a digit.

  • \:f (?:[^\[\]\:\\/<>|=+;, \t"']+) - Windows: Matches a file name part.

  • \:f (?:[^/ \t"']+) - UNIX: Matches a file name part.

  • \:h (?:[0-9A-Fa-f]+) - Matches a hex number.

  • \:i (?:[0-9]+) - Matches an integer.

  • \:n (?:(?:[0-9]+(?:\.[0-9]+|)|\.[0-9]+)(?:[Ee](?:\+|-|)[0-9]+|)) - Matches a floating number.

  • \:p (?:(?:[A-Za-z]:|)(?:\\|/|)(?:\:f(?:\\|/))*\:f) - Windows: Matches a path.

  • \:p (?:(?:/|)?:(?::f(/))*\:f) - UNIX: Matches a path.

  • \:q (?:\"[^\"]*\"|'[^']*') - Matches a quoted string.

  • \:v (?:[A-Za-z_$][A-Za-z0-9_$]*) - Matches a C variable.

  • \:w (?:[A-Za-z]+) - Matches a word.

Warning

\:f and \:p

Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed.

Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename.

The precedence of operators, from highest to lowest, is as follows:

  • +, *, ?, {}, +?, *?, ??, {}? (These operators have the same precedence.)

  • concatenation

  • |

Perl Regular Expression Examples

The table below shows examples of Perl regular expressions.

Perl Regular Expression Example

Description

^defproc

Matches lines that begin with the word defproc.

^definit$

Matches lines that only contain the word definit.

^\*name

Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *.

[\t ]

Matches tab and space characters.

[\d9\d32]

Matches tab and space characters.

[\x9\x20]

Matches tab and space characters.

p.t

Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat.

s.*?t

Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st.

for|while

Matches the strings for or while.

^\:p

Matches lines beginning with a file name.

xy+z

Matches x followed by one or more occurrences of y followed by z.

[a-z-[qw]]

Character set subtraction. Matches all English lowercase letters except q and w.

[\p{isGreek}&[\p{L}]]

Character set intersection. Matches all Unicode letters in the Greek block.

\x{6587}

Matches Unicode character with hexadecimal value 6587. Character set intersection. Matches all Unicode letters in the Greek block.

[\p{L}-[qw]]

Matches all Unicode letters except q and w.

[\p{L}]

Matches all Unicode letters.

[\p{Lul}]

Matches all Unicode uppercase and lowercase letters.

[\P{L}]

Matches all Unicode characters that are not letters.

[\p{isGreek}]

Matches all Unicode characters in the Greek block.

\x0d\x0a\x01\x02

Matches a sequence of hex binary characters.

\d13\d10\d1\d2

Matches a sequence of decimal binary characters.

SlickEdit Regular Expressions

SlickEdit regular expressions are defined in the following table.

SlickEdit Regular Expression

Definition

^

Matches beginning of line.

$

Matches end of line.

?

Matches any character except newline.

X+

Minimal match of one or more occurrences of X. See Minimal versus Maximal Matching for more information.

X#

Maximal match of one or more occurrences of X.

X*

Minimal match of zero or more occurrences of X.

X@

Maximal match of zero or more occurrences of X.

X: n1

Matches exactly n1 occurrences of X. Use () to avoid ambiguous expressions. For example a:9()1 searches for nine instance of the letter a followed by a 1.

X: n1 ,

Maximal match of at least n1 occurrences of X.

X: n1 , n2

Maximal match of at least n1 occurrences but not more than n2 occurrences of X.

X:* n1 ,

Minimal match of at least n1 occurrences of X.

X:* n1 , n2

Minimal match of at least n1 occurrences but not more than n2 occurrences of X.

~X

Search fails if expression X is matched. The expression ^~(if) matches the beginning of all lines that do not start with if.

(X)

Matches subexpression X.

{X}

Matches subexpression X and specifies a new tagged expression. See Using Tagged Expressions for more information.

{# d X}

Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9.

X|Y

Matches X or Y.

[ char-set ]

Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. Backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\0-\27] matches ASCII character codes 0..27. The expression [] matches no characters. In UNIX regular expressions, []] matches a right bracket. In both syntaxes, the expression [\]] matches a right bracket. The expression [\^] matches a caret (^) character in both syntaxes.

[~ char-set ]

Matches any character not specified by char-set. A dash (-) character may be used to specify ranges. The expression [~A-Z] matches all characters except uppercase letters.

[^ char-set ]

Same as [~char-set] above.

[ char-set1 - [ char-set2 ]]

Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w.

[ char-set1 & [ char-set2 ]]

Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127.

\x{ hhhh }

Matches up to 31-bit Unicode hexadecimal character specified by hhhh.

\p{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions.

\P{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions.

\p{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions.

\P{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions.

\x hh

Matches hexadecimal character hh where 0<=hh<=0xff.

\ ddd

Matches decimal character ddd where 0<=ddd<=255.

\g d

Defines a back reference to tagged expression number d. For example, {abc}def\g0 matches the string abcdefabc. If the tagged expression has not been set, the search fails.

\c

Specifies cursor position if match is found. If the expression xyz\c is found, the cursor is placed after the z.

\n

Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character.

\r

Matches carriage return.

\t

Matches tab character.

\b

Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word. Note that this notation previously matched a backspace character. It can still be used to match a backspace character by using it in a character set (for example, [\b]).

\B

Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word.

\Q and \E

\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching.

\f

Matches form feed character.

\od

Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing.

\om

Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om?# matches the rest of the buffer. NOTE: Test the regular expression on a very small file before using it on a large file. This option may cause the editor to use a lot of memory.

\ol

Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like ?# will not match multiple lines. This is much safer and usually faster than using the \om option.

\oi

Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside all ( ) and { } groups, after which case matching is restored to the previous case match setting. See also \oc.

\oc

Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside all ( ) and { } groups, after which case matching is restored to the previous case match setting. See also \oi.

\ char

Declares character after slash to be literal. For example, \: represents the colon character.

: char

Matches predefined expression corresponding to char. The predefined expressions are:

  • :a [A-Za-z0-9] - Matches an alphanumeric character.

  • :b ([ \t]#\) - Matches blanks - note that :b is not like the Perl/.NET \s.

  • :c [A-Za-z] - Matches an alphabetic character.

  • :d [0-9] - Matches a digit.

  • :f ([~\[\]\:\\/<>|=+;, \t"']#) - Windows: Matches a file name part.

  • :f ([~/ \t"']#) - UNIX: Matches a file name part.

  • :h ([0-9A-Fa-f]#) - Matches a hex number.

  • :i ([0-9]#) - Matches an integer.

  • :n (([0-9]#(.[0-9]#|)|.[0-9]#)([Ee](\+|-|)[0-9]#|)) - Matches a floating number.

  • :p (([A-Za-z]\:|)(\\|/|)(:f(\\|/))@:f) - Windows: Matches a path.

  • :p ((/|)(:f(/))@:f) - UNIX: Matches a path.

  • :q (\"[~\"]@\"|'[~']@') - Matches a quoted string.

  • :v ([A-Za-z_$][A-Za-z0-9_$]@) - Matches a C variable.

  • :w ([A-Za-z]#) - Matches a word.

Warning

:f and :p

Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed.

Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename.

The precedence of operators, from highest to lowest, is as follows:

  • +, #, *, @, :, :* (These operators have the same precedence.)

  • concatenation

  • |

SlickEdit Regular Expression Examples

The table below shows examples of SlickEdit regular expressions.

SlickEdit Regular Expression Example

Description

^defproc

Matches lines that begin with the word defproc.

^definit$

Matches lines that only contain the word definit.

^\:name

Matches lines that begin with the string :name. Notice that the backslash must prefix the colon character (:).

[\t ]

Matches tab and space characters.

[\9\32]

Matches tab and space characters.

[\x9\x20]

Matches tab and space characters.

p?t

Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat.

s?*t

Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st.

for|while

Matches the strings for or while.

^:p

Matches lines beginning with a file name.

xy+z

Matches x followed by one or more occurrences of y followed by z.

\x0d\x0a\x01\x02

Matches a sequence of hex binary characters.

\13\10\1\2

Matches a sequence of decimal binary characters.

UNIX Regular Expressions

UNIX regular expressions are defined in the following table.

UNIX Regular Expression

Definition

^

Matches beginning of line.

$

Matches end of line.

.

Matches any character except newline.

X+

Maximal match of one or more occurrences of X. See Minimal versus Maximal Matching.

X*

Maximal match of zero or more occurrences of X.

X?

Maximal match of zero or one occurrences of X.

X{ n1 }

Match exactly n1 occurrences of X.

X{ n1 ,}

Maximal match of at least n1 occurrences of X.

X{, n2 }

Maximal match of at least zero occurrences but not more than n2 occurrences of X.

X{ n1 , n2 }

Maximal match of at least n1 occurrences but not more than n2 occurrences of X.

X+?

Minimal match of one or more occurrences of X.

X*?

Minimal match of zero or more occurrences of X.

X??

Minimal match of zero or one occurrences of X.

X{ n1 }?

Matches exactly n1 occurrences of X.

X{ n1 ,}?

Minimal match of at least n1 occurrences of X.

X{, n2 }?

Minimal match of at least zero occurrences but not more than n2 occurrences of X.

X{ n1 , n2 }?

Minimal match of at least n1 occurrences but not more than n2 occurrences of X.

(?!X)

Search fails if expression X is matched. The expression ^(?!if) matches the beginning of all lines that do not start with if.

(?=X)

Assert, positive lookahead. Searches for subexpression X, but X is not returned as part of the match. For example, to match words ending in "ed" while excluding "ed" as part of the match, use \b[a-z]+(?=ed\b). See also (?!X).

(?>X)

Prohibit backtracking. This expression is advanced usage. It can be used to prevent the subexpression X from backtracking when using maximal (greedy) matching.

(?#text)

Comment. No text is matched in this expression; it is used for comment and documentation only.

(X)

Matches subexpression X and specifies a new tagged expression (see Using Tagged Expressions). No more tagged expressions are defined once an explicit tagged expression number is specified as shown below.

(? d X)

Matches subexpression X and specifies to use tagged expression number d where 0<=d<=9. No more tagged expressions are defined by the subexpression syntax (X) once this subexpression syntax is used. This is the best way to make sure you have enough tagged expressions.

(?:X)

Matches subexpression X but does not define a tagged expression.

X|Y

Matches X or Y.

[ char-set ]

Matches any one of the characters specified by char-set. A dash (-) character may be used to specify ranges. The expression [A-Z] matches any uppercase letter. A backslash (\) may be used inside the square brackets to define literal characters or define ASCII characters. For example, \- specifies a literal dash character. The expression [\d0-\d27] matches ASCII character codes 0..27. The expression []] matches a right bracket. In SlickEdit® Core regular expressions, [] matches no characters. In both syntaxes, the expression [\]] matches a right bracket. The expression [^] matches a caret (^) character but this does not work for SlickEdit regular expressions. In both syntaxes, [\^] matches a caret (^) character.

[^ char-set ]

Matches any character not specified by char-set. A dash (-) character may be used to specify ranges.

[ char-set1 - [ char-set2 ]]

Character set subtraction. Matches all characters in char-set1 except the characters in char-set2. The expression [^A-Z] matches all characters except uppercase letters. For example, [a-z-[qw]] matches all English lowercase letters except q and w. [\p{L}-[qw]] matches all Unicode lowercase letters except q and w.

[ char-set1 & [ char-set2 ]

Character set intersection. Matches all characters in char-set1 that are also in char-set2. For example, [\x{0}-\x{7f}&[\p{L}]] matches all letters between 0 and 127.

\x{ hhhh }

Matches up to 31-bit Unicode hexadecimal character specified by hhhh.

\p{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters in UnicodeCategorySpec. Where UnicodeCategorySpec uses the standard general categories specified by the Unicode consortium. For example, [\p{L}] matches all letters. [\p{Lu}] matches all uppercase letters. See Unicode Category Specifications for Regular Expressions.

\P{ UnicodeCategorySpec ]

(Only valid in character set) Matches characters not in UnicodeCategorySpec. For example, [\P{L}] matches all characters that are not letters. This is equivalent to [^\p{L}]. [\P{Lu}] matches all characters that are not uppercase letters. See Unicode Category Specifications for Regular Expressions.

\p{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters in UnicodeIsBlockSpec. Where UnicodeIsBlockSpec one of the standard character blocks specified by the Unicode consortium. For example, [\p{isGreek}] matches Unicode characters in the Greek block. See Unicode Character Blocks for Regular Expressions.

\P{ UnicodeIsBlockSpec ]

(Only valid in character set) Matches characters not in UnicodeIsBlockSpec. For example, [\P{isGreek}] matches all characters that are not in the Unicode Greek block. This is equivalent to [^\p{isGreek}]. See Unicode Character Blocks for Regular Expressions.

\x hh

Matches hexadecimal character hh where 0<=hh<=0xff.

\d ddd

Matches decimal character ddd where 0<=ddd<=255.

\ d

Defines a back reference to tagged expression number d. For example, {abc}def\0 matches the string abcdefabc. If the tagged expression has not been set, the search fails.

\c

Specifies cursor position if match is found. If the expression xyz\c is found the cursor is placed after the z.

\n

Matches newline character sequence. Useful for matching multi-line search strings. What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user-defined ASCII file. Use \d10 if you want to match an ASCII 10 character.

\r

Matches carriage return (ASCII 13). What this matches depends on whether the buffer is a DOS (ASCII 13,10 or just ASCII 10), UNIX (ASCII 10), Macintosh (ASCII 13), or user defined ASCII file.

\t

Matches tab character.

\b

Matches at word boundary. For example, \bre matches all occurrences of "re" that only occur at the beginning of a word.

\B

Matches all except at word boundary. For example, \Bre matches all occurrences of "re" as long as it is not at the start of a new word.

\Q and \E

\Q matches all characters as literals until \E. This is useful for longer sequences of characters without the need for the escape character. \Q does not require termination with \E, as it will continue to match characters literally until the end of the search string. \E returns to using special character tokens for matching.

\f

Matches form feed character.

\od

Matches any 2-byte DBCS character. This escape is only valid in a match set ([...\od...]). [^\od] matches any single byte character excluding end-of-line characters. When used to search Unicode text, this escape does nothing.

\om

Turns on multi-line matching. This enhances the match character set, or match any character primitives to support matching end-of-line characters. For example, \om.+ matches the rest of the buffer.

\ol

Turns off multi-line matching (default). You can still use \n to create regular expressions which match one or more lines. However, expressions like .+ will not match multiple lines. This is much safer and usually faster than using the \om option.

\oi

Ignore case. Turns off case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oc.

\oc

Case-sensitive match. Turns on case-sensitive matching in the pattern, overriding the global case setting. This modifier is localized inside the current grouping level, after which case matching is restored to the previous case match setting. See also \oi.

\ char

Declares character after slash to be literal. For example, \* represents the star character.

\: char

Matches predefined expression corresponding to char. The pre-defined expressions are:

  • \:a [A-Za-z0-9] - Matches an alphanumeric character.

  • \:c [A-Za-z] - Matches an alphabetic character.

  • \:b (?:[ \t]+) - Matches blanks.

  • \:d [0-9] - Matches a digit.

  • \:f (?:[^\[\]\:\\/<>|=+;, \t"']+) - Windows: Matches a file name part.

  • \:f (?:[^/ \t"']+) - UNIX: Matches a file name part.

  • \:h (?:[0-9A-Fa-f]+) - Matches a hex number.

  • \:i (?:[0-9]+) - Matches an integer.

  • \:n (?:(?:[0-9]+(?:\.[0-9]+|)|\.[0-9]+)(?:[Ee](?:\+|-|)[0-9]+|)) - Matches a floating number.

  • \:p (?:(?:[A-Za-z]:|)(?:\\|/|)(?:\:f(?:\\|/))*\:f) - Windows: Matches a path.

  • \:p (?:(?:/|)?:(?::f(/))*\:f) - UNIX: Matches a path.

  • \:q (?:\"[^\"]*\"|'[^']*') - Matches a quoted string.

  • \:v (?:[A-Za-z_$][A-Za-z0-9_$]*) - Matches a C variable.

  • \:w (?:[A-Za-z]+) - Matches a word.

Warning

\:f and \:p

Windows - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space characters in filenames are not allowed.

Unix - this regular expression should not be used to validate an operating system filename. The intent with this predefined regular expression is to make it useful in practice for handling filenames output from compilers and filenames in source files. For example, space, :, “, and " characters in filenames are not allowed even though the OS allows them. In the future, we may add < and > to the list of characters not allowed in a filename.

The precedence of operators, from highest to lowest, is as follows:

  • +, *, ?, {}, +?, *?, ??, {}? (These operators have the same precedence.)

  • concatenation

  • |

UNIX Regular Expression Examples

The table below shows examples of UNIX regular expressions.

UNIX Regular Expression Example

Description

^defproc

Matches lines that begin with the word defproc.

^definit$

Matches lines that only contain the word definit.

^\*name

Matches lines that begin with the string *name. Notice that the backslash must prefix the special character *.

[\t ]

Matches tab and space characters.

[\d9\d32]

Matches tab and space characters.

[\x9\x20]

Matches tab and space characters.

p.t

Matches any three-letter string starting with the letter p and ending with the letter t. Two possible matches are pot and pat.

s.*?t

Matches the letter s followed by any number of characters followed by the nearest letter t. Two possible matches are seat and st.

for|while

Matches the strings for or while.

^\:p

Matches lines beginning with a file name.

xy+z

Matches x followed by one or more occurrences of y followed by z.

[a-z-[qw]]

Character set subtraction. Matches all English lowercase letters except q and w.

[\p{isGreek}&[\p{L}]]

Character set intersection. Matches all Unicode letters in the Greek block.

\x{6587}

Matches Unicode character with hexadecimal value 6587. Character set intersection. Matches all Unicode letters in the Greek block.

[\p{L}-[qw]]

Matches all Unicode letters except q and w.

[\p{L}]

Matches all Unicode letters.

[\p{Lul}]

Matches all Unicode uppercase and lowercase letters.

[\P{L}]

Matches all Unicode characters that are not letters.

[\p{isGreek}]

Matches all Unicode characters in the Greek block.

\x0d\x0a\x01\x02

Matches a sequence of hex binary characters.

\d13\d10\d1\d2

Matches a sequence of decimal binary characters.

Wildcard Expressions

SlickEdit® Core supports *, ?, and # wildcards:

  • The asterisk (*) matches zero or more characters. For example, search for a*b to find any string that contains a lowercase letter "a" followed by a lowercase letter "b" allowing for text in between.

  • The question mark (?) matches any single character. Use multiple question marks in succession to represent that number of characters. For example, search for a???b to find any string that contains a lowercase letter "a" followed by any three characters, followed by a lowercase letter "b".

  • The pound sign (#) matches any single digit, 0-9. Use multiple pound signs in succession to represent that number of digits. For example, use ##:## to search for four-digit time-of-day values.

Unicode Categories and Character Blocks

Unicode Category Specifications for Regular Expressions

The Unicode consortium standard regular expression categories are supported. The syntax for specifying categories is:

\p{MainCategoryLetter Subcategories}

The above syntax matches the categories specified. The following syntax matches all characters not in the categories specified:

\P{MainCategoryLetter Subcategories}

The \p and \P notations can only be used inside a character set specification. MainCategoryLetter can be L, M, N, P, S, Z, or C. The valid Subcategories depend on the MainCategoryLetter specified. If no Subcategories are specified, all are assumed. For example:

  • [\p{L}] matches all Unicode letters.

  • [\p{Lul}] matches all uppercase and lowercase letters.

  • [\P{L}] matches all characters that are not letters.

The following table lists the valid subcategories for a specific main category. These character tables were generated using the file UnicodeData-3.1.0.txt found on the Unicode Consortium Web site (http://unicode.org).

Subcategory

Description

Lu

Letter, Uppercase

Ll

Letter, Lowercase

Lt

Letter, Titlecase

Lo

Letter, Other

Mn

Mark, Non-Spacing

Mc

Mark, Spacing Combining

Me

Mark, Enclosing

Nd

Number, Decimal Digit

Nl

Number, Letter

No

Number, Other

Pc

Punctuation, Connector

Pd

Punctuation, Dash

Ps

Punctuation, Open

Pe

Punctuation, Close

Pi

Punctuation, Initial quote (may behave like Ps or Pe depending on usage)

Pf

Punctuation, Final quote (may behave like Ps or Pe depending on usage)

Po

Punctuation, Other

Sm

Symbol, Math

Sc

Symbol, Currency

Sk

Symbol, Modifier

So

Symbol, Other

Zs

Separator, Space

Zl

Separator, Line

Zp

Separator, Paragraph

Cc

Other, Control

Cf

Other, Format

Cs

Other, Surrogate

Co

Other, Private Use

Cn

Other, Not Assigned (no characters in the file have this property)

Unicode Character Blocks for Regular Expressions

The Unicode consortium standard regular expression block categories are supported. The syntax for specifying a character block is:

\p{Is BlockName}

The above syntax matches the characters in the block specified. The following syntax matches all characters not in the block specified:

\P{Is BlockName }

The \p and \P notations may only be used inside a character set specification. For example, [\p{isBasicLatin}] matches all characters in the Greek block. [\P{isBasicLatin}] matches all characters that are not in the Greek block.

The following table lists the non-standard valid character block names. These character tables were generated from XML standards found at the World Wide Web Consortium Web site).

Block Name

Description

XMLNameStartChar

All characters that are valid for the start of an XML tag name.

XMLNameChar

All characters that are valid in an XML tag name.

The following table lists the valid character block names. These character tables were generated using the blocks.txt file found on the Unicode Consortium Web site (http://unicode.org).

Range

Block Name

0000..007F

BasicLatin

0080..00FF

Latin-1Supplement

0100..017F

LatinExtended-A

0180..024F

LatinExtended-B

0250..02AF

IPAExtensions

02B0..02FF

SpacingModifierLetters

0300..036F

CombiningDiacriticalMarks

0370..03FF

Greek

0400..04FF

Cyrillic

0530..058F

Armenian

0590..05FF

Hebrew

0600..06FF

Arabic

0700..074F

Syriac

0780..07BF

Thaana

0900..097F

Devanagari

0980..09FF

Bengali

0A00..0A7F

Gurmukhi

0A80..0AFF

Gujarati

0B00..0B7F

Oriya

0B80..0BFF

Tamil

0C00..0C7F

Telugu

0C80..0CFF

Kannada

0D00..0D7F

Malayalam

0D80..0DFF

Sinhala

0E00..0E7F

Thai

0E80..0EFF

Lao

0F00..0FFF

Tibetan

1000..109F

Myanmar

10A0..10FF

Georgian

1100..11FF

HangulJamo

1200..137F

Ethiopic

13A0..13FF

Cherokee

1400..167F

UnifiedCanadianAboriginalSyllabics

1680..169F

Ogham

16A0..16FF

Runic

1780..17FF

Khmer

1800..18AF

Mongolian

1E00..1EFF

LatinExtendedAdditional

1F00..1FFF

GreekExtended

2000..206F

GeneralPunctuation

2070..209F

SuperscriptsandSubscripts

20A0..20CF

CurrencySymbols

20D0..20FF

CombiningMarksforSymbols

2100..214F

LetterlikeSymbols

2150..218F

NumberForms

2190..21FF

Arrows

2200..22FF

MathematicalOperators

2300..23FF

MiscellaneousTechnical

2400..243F

ControlPictures

2440..245F

OpticalCharacterRecognition

2460..24FF

EnclosedAlphanumerics

2500..257F

BoxDrawing

2580..259F

BlockElements

25A0..25FF

GeometricShapes

2600..26FF

MiscellaneousSymbols

2700..27BF

Dingbats

2800..28FF

BraillePatterns

2E80..2EFF

CJKRadicalsSupplement

2F00..2FDF

KangxiRadicals

2FF0..2FFF

IdeographicDescriptionCharacters

3000..303F

CJKSymbolsandPunctuation

3040..309F

Hiragana

30A0..30FF

Katakana

3100..312F

Bopomofo

3130..318F

HangulCompatibilityJamo

3190..319F

Kanbun

31A0..31BF

BopomofoExtended

3200..32FF

EnclosedCJKLettersandMonths

3300..33FF

CJKCompatibility

3400..4DB5

CJKUnifiedIdeographsExtensionA

4E00..9FFF

CJKUnifiedIdeographs

A000..A48F

YiSyllables

A490..A4CF

YiRadicals

AC00..D7A3

HangulSyllables

D800..DB7F

HighSurrogates

DB80..DBFF

HighPrivateUseSurrogates

DC00..DFFF

LowSurrogates

E000..F8FF

PrivateUse

F900..FAFF

CJKCompatibilityIdeographs

FB00..FB4F

AlphabeticPresentationForms

FB50..FDFF

ArabicPresentationForms-A

FE20..FE2F

CombiningHalfMarks

FE30..FE4F

CJKCompatibilityForms

FE50..FE6F

SmallFormVariants

FE70..FEFE

ArabicPresentationForms-B

FEFF..FEFF

Specials

FF00..FFEF

HalfwidthandFullwidthForms

FFF0..FFFD

Specials

10300..1032F

OldItalic

10330..1034F

Gothic

10400..1044F

Deseret

1D000..1D0FF

ByzantineMusicalSymbols

1D100..1D1FF

MusicalSymbols

1D400..1D7FF

MathematicalAlphanumericSymbols

20000..2A6D6

CJKUnifiedIdeographsExtensionB

2F800..2FA1F

CJKCompatibilityIdeographsSupplement

E0000..E007F

Tags