Class SequenceMatchRules


  • public class SequenceMatchRules
    extends Object
    Rules for matching sequences using regular expressions.

    There are 2 types of rules:

    1. Assignment rules which assign a value to a variable for later use.
    2. Extraction rules which specifies how regular expression patterns are to be matched against text, which matched text expressions are to extracted, and what value to assign to the matched expression.
    NOTE: # or // can be used to indicates one-line comments.

    Assignment Rules are used to assign values to variables. The basic format is: variable = value.

    Variable Names:

    • Variable names should follow the pattern [A-Za-z_][A-Za-z0-9_]*
    • Variable names for use in regular expressions (to be expanded later) must start with $

    Value Types:

    Value Types
    TypeFormatExampleDescription
    BOOLEANTRUE | FALSETRUE
    STRING"...""red"
    INTEGER[+-]\d+1500
    LONG[+-]\d+L1500000000000L
    DOUBLE[+-]\d*\.\d+6.98
    REGEX/...//[Aa]pril/ String regular expression Pattern
    TOKENS_REGEX( [...] [...] ... ) ( /up/ /to/ /4/ /months/ ) Tokens regular expression TokenSequencePattern
    LIST( [item1] , [item2], ... )("red", "blue", "yellow" )

    Some typical uses and examples for assignment rules include:

    1. Assignment of value to variables for use in later rules
    2. Binding of text key to annotation key (as Class).
            tokens = { type: "CLASS", value: "edu.stanford.nlp.ling.CoreAnnotations$TokensAnnotation" }
          
    3. Defining regular expressions macros to be embedded in other regular expressions
            $SEASON = "/spring|summer|fall|autumn|winter/"
            $NUM = ( [ { numcomptype:NUMBER } ] )
          
    4. Setting default environment variables. Rules are applied with respect to an environment (Env), which can be accessed using the variable ENV. Members of the Environment can be set as needed.
            # Set default parameters to be used when reading rules
            ENV.defaults["ruleType"] = "tokens"
            # Set default string pattern flags (to case-insensitive)
            ENV.defaultStringPatternFlags = 2
            # Specifies that the result should go into the tokens  key (as defined above).
            ENV.defaultResultAnnotationKey = tokens
          
    5. Defining options

    Predefined values are:

    Predefined values
    VariableTypeDescription
    ENVEnvThe environment with respect to which the rules are applied.
    TRUEBOOLEANThe Boolean value true.
    FALSEBOOLEANThe Boolean value false.
    NILThe null value.
    tagsClassThe annotation key Tags.TagsAnnotation.

    Extraction Rules specifies how regular expression patterns are to be matched against text. See CoreMapExpressionExtractor for more information on the types of the rules, and in what sequence the rules are applied. A basic rule can be specified using the following template:

       {
         # Type of the rule
         ruleType: "tokens" | "text" | "composite" | "filter",
         # Pattern to match against
         pattern: ( <TokenSequencePattern> ) | /<TextPattern>/,
         # Resulting value to go into the resulting annotation
         result: ...
    
         # More fields following...
       }
     
    Example:
       {
         ruleType: "tokens",
         pattern: ( /one/ ),
         result: 1
       }
     

    Extraction rule fields (most fields are optional):

    Extraction rule fields
    FieldValuesExampleDescription
    ruleType"tokens" | "text" | "composite" | "filter" tokensType of the rule (required).
    pattern<Token Sequence Pattern> = (...) | <Text Pattern> = /.../ ( /winter/ /of/ $YEAR )Pattern to match against. See TokenSequencePattern and Pattern for how to specify patterns over tokens and strings (required).
    action<Action List> = (...) ( Annotate($0, ner, "DATE") )List of actions to apply when the pattern is triggered. Each action is a TokensRegex Expression
    result<Expression> Resulting value to go into the resulting annotation. See Expressions for how to specify the result.
    nameSTRING Name to identify the extraction rule.
    stageINTEGER Stage at which the rule is to be applied. Rules are grouped in stages, which are applied from lowest to highest.
    activeBoolean Whether this rule is enabled (active) or not (default true).
    priorityDOUBLE Priority of rule. Within a stage, matches from higher priority rules are preferred.
    weightDOUBLE Weight of rule (not currently used).
    overCLASS Annotation field to check pattern against.
    matchFindTypeFIND_NONOVERLAPPING | FIND_ALL Whether to find all matched expression or just the nonoverlapping ones (default FIND_NONOVERLAPPING).
    matchWithResultsBoolean Whether results of the matches should be returned (default false). Set to true to access captured groups of embedded regular expressions.
    matchedExpressionGroupInteger 2What group should be treated as the matched expression group (default 0).
    Author:
    Angel Chang
    See Also:
    CoreMapExpressionExtractor, TokenSequencePattern