Class SequenceMatchRules
- java.lang.Object
-
- edu.stanford.nlp.ling.tokensregex.SequenceMatchRules
-
public class SequenceMatchRules extends Object
Rules for matching sequences using regular expressions.There are 2 types of rules:
- Assignment rules which assign a value to a variable for later use.
- Extraction rules which specifies how regular expression patterns are to be matched against text, which matched text expressions are to extracted, and what value to assign to the matched expression.
#or//can be used to indicates one-line comments.Assignment Rules are used to assign values to variables. The basic format is:
variable = value.Variable Names:
- Variable names should follow the pattern [A-Za-z_][A-Za-z0-9_]*
- Variable names for use in regular expressions (to be expanded later) must start with
$
Value Types:
Value Types Type Format Example Description BOOLEANTRUE | FALSETRUESTRING"...""red"INTEGER[+-]\d+1500LONG[+-]\d+L1500000000000LDOUBLE[+-]\d*\.\d+6.98REGEX/...//[Aa]pril/String regular expression PatternTOKENS_REGEX( [...] [...] ... )( /up/ /to/ /4/ /months/ )Tokens regular expression TokenSequencePatternLIST( [item1] , [item2], ... )("red", "blue", "yellow" )Some typical uses and examples for assignment rules include:
- Assignment of value to variables for use in later rules
- Binding of text key to annotation key (as
Class).tokens = { type: "CLASS", value: "edu.stanford.nlp.ling.CoreAnnotations$TokensAnnotation" } - Defining regular expressions macros to be embedded in other regular expressions
$SEASON = "/spring|summer|fall|autumn|winter/" $NUM = ( [ { numcomptype:NUMBER } ] ) - Setting default environment variables.
Rules are applied with respect to an environment (
Env), which can be accessed using the variableENV. Members of the Environment can be set as needed.# Set default parameters to be used when reading rules ENV.defaults["ruleType"] = "tokens" # Set default string pattern flags (to case-insensitive) ENV.defaultStringPatternFlags = 2 # Specifies that the result should go into thetokenskey (as defined above). ENV.defaultResultAnnotationKey = tokens - Defining options
Predefined values are:
Predefined values Variable Type Description ENVEnvThe environment with respect to which the rules are applied. TRUEBOOLEANThe Booleanvaluetrue.FALSEBOOLEANThe Booleanvaluefalse.NILThe nullvalue.tagsClassThe annotation key Tags.TagsAnnotation.Extraction Rules specifies how regular expression patterns are to be matched against text. See
CoreMapExpressionExtractorfor more information on the types of the rules, and in what sequence the rules are applied. A basic rule can be specified using the following template:{ # Type of the rule ruleType: "tokens" | "text" | "composite" | "filter", # Pattern to match against pattern: ( <TokenSequencePattern> ) | /<TextPattern>/, # Resulting value to go into the resulting annotation result: ... # More fields following... }Example:{ ruleType: "tokens", pattern: ( /one/ ), result: 1 }Extraction rule fields (most fields are optional):
Extraction rule fields Field Values Example Description ruleType"tokens" | "text" | "composite" | "filter"tokensType of the rule (required). pattern<Token Sequence Pattern> = (...) | <Text Pattern> = /.../( /winter/ /of/ $YEAR )Pattern to match against. See TokenSequencePatternandPatternfor how to specify patterns over tokens and strings (required).action<Action List> = (...)( Annotate($0, ner, "DATE") )List of actions to apply when the pattern is triggered. Each action is a TokensRegex Expressionresult<Expression>Resulting value to go into the resulting annotation. See Expressionsfor how to specify the result.nameSTRINGName to identify the extraction rule. stageINTEGERStage at which the rule is to be applied. Rules are grouped in stages, which are applied from lowest to highest. activeBooleanWhether this rule is enabled (active) or not (default true). priorityDOUBLEPriority of rule. Within a stage, matches from higher priority rules are preferred. weightDOUBLEWeight of rule (not currently used). overCLASSAnnotation field to check pattern against. matchFindTypeFIND_NONOVERLAPPING | FIND_ALLWhether to find all matched expression or just the nonoverlapping ones (default FIND_NONOVERLAPPING).matchWithResultsBooleanWhether results of the matches should be returned (default false). Set to true to access captured groups of embedded regular expressions. matchedExpressionGroupInteger2What group should be treated as the matched expression group (default 0). - Author:
- Angel Chang
- See Also:
CoreMapExpressionExtractor,TokenSequencePattern
-
-
Nested Class Summary
-
Field Summary
-
Method Summary
-
-
-
Field Detail
-
COMPOSITE_RULE_TYPE
public static final String COMPOSITE_RULE_TYPE
- See Also:
- Constant Field Values
-
TOKEN_PATTERN_RULE_TYPE
public static final String TOKEN_PATTERN_RULE_TYPE
- See Also:
- Constant Field Values
-
TEXT_PATTERN_RULE_TYPE
public static final String TEXT_PATTERN_RULE_TYPE
- See Also:
- Constant Field Values
-
FILTER_RULE_TYPE
public static final String FILTER_RULE_TYPE
- See Also:
- Constant Field Values
-
TOKEN_PATTERN_EXTRACT_RULE_CREATOR
public static final SequenceMatchRules.TokenPatternExtractRuleCreator TOKEN_PATTERN_EXTRACT_RULE_CREATOR
-
COMPOSITE_EXTRACT_RULE_CREATOR
public static final SequenceMatchRules.CompositeExtractRuleCreator COMPOSITE_EXTRACT_RULE_CREATOR
-
TEXT_PATTERN_EXTRACT_RULE_CREATOR
public static final SequenceMatchRules.TextPatternExtractRuleCreator TEXT_PATTERN_EXTRACT_RULE_CREATOR
-
MULTI_TOKEN_PATTERN_EXTRACT_RULE_CREATOR
public static final SequenceMatchRules.MultiTokenPatternExtractRuleCreator MULTI_TOKEN_PATTERN_EXTRACT_RULE_CREATOR
-
DEFAULT_EXTRACT_RULE_CREATOR
public static final SequenceMatchRules.AnnotationExtractRuleCreator DEFAULT_EXTRACT_RULE_CREATOR
-
-
Method Detail
-
createAssignmentRule
public static SequenceMatchRules.AssignmentRule createAssignmentRule(Env env, AssignableExpression var, Expression result)
-
createRule
public static SequenceMatchRules.Rule createRule(Env env, Expressions.CompositeValue cv)
-
createExtractionRule
protected static SequenceMatchRules.AnnotationExtractRule createExtractionRule(Env env, Map<String,Object> attributes)
-
createExtractionRule
public static SequenceMatchRules.AnnotationExtractRule createExtractionRule(Env env, String ruleType, Object pattern, Expression result)
-
createTokenPatternRule
public static SequenceMatchRules.AnnotationExtractRule createTokenPatternRule(Env env, SequencePattern.PatternExpr expr, Expression result)
-
createTextPatternRule
public static SequenceMatchRules.AnnotationExtractRule createTextPatternRule(Env env, String expr, Expression result)
-
createMultiTokenPatternRule
public static SequenceMatchRules.AnnotationExtractRule createMultiTokenPatternRule(Env env, SequenceMatchRules.AnnotationExtractRule template, List<TokenSequencePattern> patterns)
-
createAnnotationExtractor
public static MatchedExpression.SingleAnnotationExtractor createAnnotationExtractor(Env env, SequenceMatchRules.AnnotationExtractRule r)
-
-