Class SpanishTokenizer.SpanishTokenizerFactory<T extends HasWord>

    • Field Detail

      • lexerProperties

        protected Properties lexerProperties
      • splitCompoundOption

        protected boolean splitCompoundOption
      • splitVerbOption

        protected boolean splitVerbOption
      • splitContractionOption

        protected boolean splitContractionOption
    • Method Detail

      • newSpanishTokenizerFactory

        public static <T extends HasWord> SpanishTokenizer.SpanishTokenizerFactory<T> newSpanishTokenizerFactory​(LexedTokenFactory<T> factory,
                                                                                                                 String options)
        Constructs a new SpanishTokenizer that returns T objects and uses the options passed in.
        Parameters:
        options - a String of options, separated by commas
        factory - a factory for the token type that the tokenizer will return
        Returns:
        A TokenizerFactory that returns the right token types
      • setOptions

        public void setOptions​(String options)
        Set underlying tokenizer options.
        Specified by:
        setOptions in interface TokenizerFactory<T extends HasWord>
        Parameters:
        options - A comma-separated list of options
      • getTokenizer

        public Tokenizer<T> getTokenizer​(Reader r,
                                         String extraOptions)
        Description copied from interface: TokenizerFactory
        Get a tokenizer for this reader.
        Specified by:
        getTokenizer in interface TokenizerFactory<T extends HasWord>
        Parameters:
        r - A Reader (which is assumed to already by buffered, if appropriate)
        extraOptions - Options for how this tokenizer should behave
        Returns:
        A Tokenizer