Class CMMClassifier<IN extends CoreLabel>

  • All Implemented Interfaces:
    Function<String,​String>

    public class CMMClassifier<IN extends CoreLabel>
    extends AbstractSequenceClassifier<IN>
    Does Sequence Classification using a Conditional Markov Model. It could be used for other purposes, but the provided features are aimed at doing Named Entity Recognition. The code has functionality for different document encodings, but when using the standard ColumnDocumentReader, input files are expected to be one word per line with the columns indicating things like the word, POS, chunk, and class. Typical usage For running a trained model with a provided serialized classifier: java -server -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -loadClassifier conll.ner.gz -textFile samplesentences.txt When specifying all parameters in a properties file (train, test, or runtime): java -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -prop propFile To train and test a model from the command line: java -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -trainFile trainFile -testFile testFile -goodCoNLL &gt; output Features are defined by a FeatureFactory; the FeatureFactory which is used by default is NERFeatureFactory, and you should look there for feature templates. Features are specified either by a Properties file (which is the recommended method) or on the command line. The features are read into a SeqClassifierFlags object, which the user need not know much about, unless one wishes to add new features. CMMClassifier may also be used programmatically. When creating a new instance, you must specify a properties file. The other way to get a CMMClassifier is to deserialize one via getClassifier(String), which returns a deserialized classifier. You may then tag sentences using either the assorted test or testSentence methods.
    Author:
    Dan Klein, Jenny Finkel, Christopher Manning, Shipra Dingare, Huy Nguyen, Sarah Spikes (sdspikes@cs.stanford.edu) - cleanup and filling in types