Extended Regular Expressions

As intrepreted by Dennis German based on opengroup.org/onlinepubs The extended regular expression (ERE) notation and construction rules shall apply to utilities defined as using extended regular expressions; any exceptions to the following rules are noted in the descriptions of the specific utilities using EREs. 9.4.1 EREs Matching a Single Character or Collating Element An ERE ordinary character, a special character preceded by a or a shall match a single character. A bracket expression shall match a single character or a single collating element. An ERE matching a single character enclosed in parentheses shall match the same as the ERE without parentheses would have matched. 9.4.2 ERE Ordinary Characters An ordinary character is an ERE that matches itself. An ordinary character is any character in the supported character set, except for the ERE special characters listed in ERE Special Characters. The interpretation of an ordinary character preceded by an unescaped ( '\\' ) is undefined, except in the context of a bracket expression (see ERE Bracket Expression). 9.4.3 ERE Special Characters An ERE special character has special properties in certain contexts. Outside those contexts, or when preceded by a , such a character shall be an ERE that matches the special character itself. The extended regular expression special characters and the contexts in which they shall have their special meaning are as follows: .[\( The , , , and shall be special except when used in a bracket expression (see RE Bracket Expression). Outside a bracket expression, a immediately followed by a produces undefined results. A that is unescaped and is not part of a bracket expression also produces undefined results. ) The shall be special when matched with a preceding , both outside a bracket expression. *+?{ The , , , and shall be special except when used in a bracket expression (see RE Bracket Expression). Any of the following uses produce undefined results: If these characters appear first in an ERE, or immediately following an unescaped , , , or If a is not part of a valid interval expression (see EREs Matching Multiple Characters) | The is special except when used in a bracket expression (see RE Bracket Expression). A appearing first or last in an ERE, or immediately following a or a , or immediately preceding a , produces undefined results. ^ The shall be special when used as an anchor (see ERE Expression Anchoring). The shall signify a non-matching list expression when it occurs first in a list, immediately following a (see RE Bracket Expression). $ The shall be special when used as an anchor. 9.4.4 Periods in EREs A ( '.' ), when used outside a bracket expression, is an ERE that shall match any character in the supported character set except NUL. 9.4.5 ERE Bracket Expression The rules for ERE Bracket Expressions are the same as for Basic Regular Expressions; see RE Bracket Expression. 9.4.6 EREs Matching Multiple Characters The following rules shall be used to construct EREs matching multiple characters from EREs matching a single character: A concatenation of EREs shall match the concatenation of the character sequences matched by each component of the ERE. A concatenation of EREs enclosed in parentheses shall match whatever the concatenation without the parentheses matches. For example, both the ERE "cd" and the ERE "(cd)" are matched by the third and fourth character of the string "abcdefabcdef". When an ERE matching a single character or an ERE enclosed in parentheses is followed by the special character ( '+' ), together with that it shall match what one or more consecutive occurrences of the ERE would match. For example, the ERE "b+(bc)" matches the fourth to seventh characters in the string "acabbbcde". And, "[ab]+" and "[ab][ab]*" are equivalent. When an ERE matching a single character or an ERE enclosed in parentheses is followed by the special character ( '*' ), together with that it shall match what zero or more consecutive occurrences of the ERE would match. For example, the ERE "b*c" matches the first character in the string "cabbbcde", and the ERE "b*cd" matches the third to seventh characters in the string "cabbbcdebbbbbbcdbc". And, "[ab]*" and "[ab][ab]" are equivalent when matching the string "ab". When an ERE matching a single character or an ERE enclosed in parentheses is followed by the special character ( '?' ), together with that it shall match what zero or one consecutive occurrences of the ERE would match. For example, the ERE "b?c" matches the second character in the string "acabbbcde". When an ERE matching a single character or an ERE enclosed in parentheses is followed by an interval expression of the format "{m}", "{m,}", or "{m,n}", together with that interval expression it shall match what repeated consecutive occurrences of the ERE would match. The values of m and n are decimal integers in the range 0 <= m<= n<= {RE_DUP_MAX}, where m specifies the exact or minimum number of occurrences and n specifies the maximum number of occurrences. The expression "{m}" matches exactly m occurrences of the preceding ERE, "{m,}" matches at least m occurrences, and "{m,n}" matches any number of occurrences between m and n, inclusive. For example, in the string "abababccccccd" the ERE "c{3}" is matched by characters seven to nine and the ERE "(ab){2,}" is matched by characters one to six. The behavior of multiple adjacent duplication symbols ( '+', '*', '?', and intervals) produces undefined results. An ERE matching a single character repeated by an '*', '?', or an interval expression shall not match a null expression unless this is the only match for the repetition or it is necessary to satisfy the exact or minimum number of occurrences for the interval expression. 9.4.7 ERE Alternation Two EREs separated by the special character ( '|' ) shall match a string that is matched by either. For example, the ERE "a((bc)|d)" matches the string "abc" and the string "ad". Single characters, or expressions matching single characters, separated by the and enclosed in parentheses, shall be treated as an ERE matching a single character. 9.4.8 ERE Precedence The order of precedence shall be as shown in the following table: ERE Precedence (from high to low) Collation-related bracket symbols [==] [::] [..] Escaped characters \ Bracket expression [] Grouping () Single-character-ERE duplication * + ? {m,n} Concatenation Anchoring ^ $ Alternation | For example, the ERE "abba|cde" matches either the string "abba" or the string "cde" (rather than the string "abbade" or "abbcde", because concatenation has a higher order of precedence than alternation). 9.4.9 ERE Expression Anchoring An ERE can be limited to matching expressions that begin or end a string; this is called "anchoring". The and special characters shall be considered ERE anchors when used anywhere outside a bracket expression. This shall have the following effects: A ( '^' ) outside a bracket expression shall anchor the expression or subexpression it begins to the beginning of a string; such an expression or subexpression can match only a sequence starting at the first character of a string. For example, the EREs "^ab" and "(^ab)" match "ab" in the string "abcdef", but fail to match in the string "cdefab", and the ERE "a^b" is valid, but can never match because the 'a' prevents the expression "^b" from matching starting at the first character. A ( '$' ) outside a bracket expression shall anchor the expression or subexpression it ends to the end of a string; such an expression or subexpression can match only a sequence ending at the last character of a string. For example, the EREs "ef$" and "(ef$)" match "ef" in the string "abcdef", but fail to match in the string "cdefab", and the ERE "e$f" is valid, but can never match because the 'f' prevents the expression "e$" from matching ending at the last character. 9.5 Regular Expression Grammar Grammars describing the syntax of both basic and extended regular expressions are presented in this section. The grammar takes precedence over the text. See XCU Grammar Conventions. 9.5.1 BRE/ERE Grammar Lexical Conventions The lexical conventions for regular expressions are as described in this section. Except as noted, the longest possible token or delimiter beginning at a given point is recognized. The following tokens are processed (in addition to those string constants shown in the grammar): COLL_ELEM_SINGLE Any single-character collating element, unless it is a META_CHAR. COLL_ELEM_MULTI Any multi-character collating element. BACKREF Applicable only to basic regular expressions. The character string consisting of a character followed by a single-digit numeral, '1' to '9'. DUP_COUNT Represents a numeric constant. It shall be an integer in the range 0 <= DUP_COUNT <= {RE_DUP_MAX}. This token is only recognized when the context of the grammar requires it. At all other times, digits not preceded by a character are treated as ORD_CHAR. META_CHAR One of the characters: ^ When found first in a bracket expression - When found anywhere but first (after an initial '^', if any) or last in a bracket expression, or as the ending range point in a range expression ] When found anywhere but first (after an initial '^', if any) in a bracket expression L_ANCHOR Applicable only to basic regular expressions. The character '^' when it appears as the first character of a basic regular expression and when not QUOTED_CHAR. The '^' may be recognized as an anchor elsewhere; see BRE Expression Anchoring. ORD_CHAR A character, other than one of the special characters in SPEC_CHAR. QUOTED_CHAR In a BRE, one of the character sequences: \^ \. \* \[ \$ \\ In an ERE, one of the character sequences: \^ \. \[ \$ \( \) \| \* \+ \? \{ \\ R_ANCHOR (Applicable only to basic regular expressions.) The character '$' when it appears as the last character of a basic regular expression and when not QUOTED_CHAR. The '$' may be recognized as an anchor elsewhere; see BRE Expression Anchoring. SPEC_CHAR For basic regular expressions, one of the following special characters: . Anywhere outside bracket expressions \ Anywhere outside bracket expressions [ Anywhere outside bracket expressions ^ When used as an anchor (see BRE Expression Anchoring) $ When used as an anchor * Anywhere except first in an entire RE, anywhere in a bracket expression, directly following "\(", directly following an anchoring '^' For extended regular expressions, shall be one of the following special characters found anywhere outside bracket expressions: ^ . [ $ ( ) | * + ? { \ (The close-parenthesis shall be considered special in this context only if matched with a preceding open-parenthesis.) 9.5.2 RE and Bracket Expression Grammar This section presents the grammar for basic regular expressions, including the bracket expression grammar that is common to both BREs and EREs. %token ORD_CHAR QUOTED_CHAR DUP_COUNT %token BACKREF L_ANCHOR R_ANCHOR %token Back_open_paren Back_close_paren /* '\(' '\)' */ %token Back_open_brace Back_close_brace /* '\{' '\}' */ /* The following tokens are for the Bracket Expression grammar common to both REs and EREs. */ %token COLL_ELEM_SINGLE COLL_ELEM_MULTI META_CHAR %token Open_equal Equal_close Open_dot Dot_close Open_colon Colon_close /* '[=' '=]' '[.' '.]' '[:' ':]' */ %token class_name /* class_name is a keyword to the LC_CTYPE locale category */ /* (representing a character class) in the current locale */ /* and is only recognized between [: and :] */ %start basic_reg_exp %% /* -------------------------------------------- Basic Regular Expression -------------------------------------------- */ basic_reg_exp : RE_expression | L_ANCHOR | R_ANCHOR | L_ANCHOR R_ANCHOR | L_ANCHOR RE_expression | RE_expression R_ANCHOR | L_ANCHOR RE_expression R_ANCHOR ; RE_expression : simple_RE | RE_expression simple_RE ; simple_RE : nondupl_RE | nondupl_RE RE_dupl_symbol ; nondupl_RE : one_char_or_coll_elem_RE | Back_open_paren RE_expression Back_close_paren | BACKREF ; one_char_or_coll_elem_RE : ORD_CHAR | QUOTED_CHAR | '.' | bracket_expression ; RE_dupl_symbol : '*' | Back_open_brace DUP_COUNT Back_close_brace | Back_open_brace DUP_COUNT ',' Back_close_brace | Back_open_brace DUP_COUNT ',' DUP_COUNT Back_close_brace ; /* -------------------------------------------- Bracket Expression ------------------------------------------- */ bracket_expression : '[' matching_list ']' | '[' nonmatching_list ']' ; matching_list : bracket_list ; nonmatching_list : '^' bracket_list ; bracket_list : follow_list | follow_list '-' ; follow_list : expression_term | follow_list expression_term ; expression_term : single_expression | range_expression ; single_expression : end_range | character_class | equivalence_class ; range_expression : start_range end_range | start_range '-' ; start_range : end_range '-' ; end_range : COLL_ELEM_SINGLE | collating_symbol ; collating_symbol : Open_dot COLL_ELEM_SINGLE Dot_close | Open_dot COLL_ELEM_MULTI Dot_close | Open_dot META_CHAR Dot_close ; equivalence_class : Open_equal COLL_ELEM_SINGLE Equal_close | Open_equal COLL_ELEM_MULTI Equal_close ; character_class : Open_colon class_name Colon_close ; The BRE grammar does not permit L_ANCHOR or R_ANCHOR inside "\(" and "\)" (which implies that '^' and '$' are ordinary characters). This reflects the semantic limits on the application, as noted in BRE Expression Anchoring. Implementations are permitted to extend the language to interpret '^' and '$' as anchors in these locations, and as such, conforming applications cannot use unescaped '^' and '$' in positions inside "\(" and "\)" that might be interpreted as anchors. 9.5.3 ERE Grammar This section presents the grammar for extended regular expressions, excluding the bracket expression grammar. Note: The bracket expression grammar and the associated %token lines are identical between BREs and EREs. It has been omitted from the ERE section to avoid unnecessary editorial duplication. %token ORD_CHAR QUOTED_CHAR DUP_COUNT %start extended_reg_exp %% /* -------------------------------------------- Extended Regular Expression -------------------------------------------- */ extended_reg_exp : ERE_branch | extended_reg_exp '|' ERE_branch ; ERE_branch : ERE_expression | ERE_branch ERE_expression ; ERE_expression : one_char_or_coll_elem_ERE | '^' | '$' | '(' extended_reg_exp ')' | ERE_expression ERE_dupl_symbol ; one_char_or_coll_elem_ERE : ORD_CHAR | QUOTED_CHAR | '.' | bracket_expression ; ERE_dupl_symbol : '*' | '+' | '?' | '{' DUP_COUNT '}' | '{' DUP_COUNT ',' '}' | '{' DUP_COUNT ',' DUP_COUNT '}' ; The ERE grammar does not permit several constructs that previous sections specify as having undefined results. Additionally, there are some constructs which the grammar permits but which still give undefined results: ORD_CHAR preceded by an unescaped character One or more ERE_dupl_symbols appearing first in an ERE, or immediately following '|', '^', '(', or '$' '{' not part of a valid ERE_dupl_symbol '|' appearing first or last in an ERE, or immediately following '|' or '(', or immediately preceding ')' Implementations are permitted to extend the language to allow these. Strictly Conforming applications cannot use such constructs.

opengroup.org/onlinepubs