Package edu.stanford.nlp.ie.ner
Class CMMClassifier<IN extends CoreLabel>
- java.lang.Object
-
- edu.stanford.nlp.ie.AbstractSequenceClassifier<IN>
-
- edu.stanford.nlp.ie.ner.CMMClassifier<IN>
-
public class CMMClassifier<IN extends CoreLabel> extends AbstractSequenceClassifier<IN>
Does Sequence Classification using a Conditional Markov Model. It could be used for other purposes, but the provided features are aimed at doing Named Entity Recognition. The code has functionality for different document encodings, but when using the standardColumnDocumentReader, input files are expected to be one word per line with the columns indicating things like the word, POS, chunk, and class. Typical usage For running a trained model with a provided serialized classifier:java -server -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -loadClassifier conll.ner.gz -textFile samplesentences.txtWhen specifying all parameters in a properties file (train, test, or runtime):java -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -prop propFileTo train and test a model from the command line:java -mx1000m edu.stanford.nlp.ie.ner.CMMClassifier -trainFile trainFile -testFile testFile -goodCoNLL > outputFeatures are defined by aFeatureFactory; theFeatureFactorywhich is used by default isNERFeatureFactory, and you should look there for feature templates. Features are specified either by a Properties file (which is the recommended method) or on the command line. The features are read into aSeqClassifierFlagsobject, which the user need not know much about, unless one wishes to add new features. CMMClassifier may also be used programmatically. When creating a new instance, you must specify a properties file. The other way to get a CMMClassifier is to deserialize one viagetClassifier(String), which returns a deserialized classifier. You may then tag sentences using either the assortedtestortestSentencemethods.- Author:
- Dan Klein, Jenny Finkel, Christopher Manning, Shipra Dingare, Huy Nguyen, Sarah Spikes (sdspikes@cs.stanford.edu) - cleanup and filling in types
-
-
Field Summary
Fields Modifier and Type Field Description static StringDEFAULT_CLASSIFIERDefault place to look in Jar file for classifier.-
Fields inherited from class edu.stanford.nlp.ie.AbstractSequenceClassifier
classIndex, featureFactories, flags, knownLCWords, pad, windowSize
-
-
Constructor Summary
Constructors Modifier Constructor Description protectedCMMClassifier()CMMClassifier(SeqClassifierFlags flags)CMMClassifier(Properties props)
-
Method Summary
All Methods Static Methods Instance Methods Concrete Methods Modifier and Type Method Description voidadapt(ObjectBank<List<IN>> featureLabels, Dataset<String,String> trainDataset)voidadapt(String filename, Dataset<String,String> trainDataset, DocumentReaderAndWriter<IN> readerWriter)List<IN>classify(List<IN> document)List<IN>classifyWithGlobalInformation(List<IN> tokenSeq, CoreMap doc, CoreMap sent)protected StringclassOf(List<IN> lineInfos, int pos)Returns the most likely class for the word at the given position.Dataset<String,String>getBiasedDataset(ObjectBank<List<IN>> data, Index<String> featureIndex, Index<String> classIndex)static CMMClassifier<? extends CoreLabel>getClassifier(File file)static CMMClassifier<? extends CoreLabel>getClassifier(InputStream in)static <INN extends CoreMap>
CMMClassifier<? extends CoreLabel>getClassifier(ObjectInputStream ois)static <INN extends CoreMap>
CMMClassifier<? extends CoreLabel>getClassifier(ObjectInputStream ois, Properties props)static CMMClassifier<? extends CoreLabel>getClassifier(String loadPath)static CMMClassifier<? extends CoreLabel>getClassifierNoExceptions(File file)static CMMClassifier<? extends CoreLabel>getClassifierNoExceptions(InputStream in)static CMMClassifier<CoreLabel>getClassifierNoExceptions(String loadPath)Dataset<String,String>getDataset(Dataset<String,String> oldData, Index<String> goodFeatures)Build a Dataset from some data.Dataset<String,String>getDataset(ObjectBank<List<IN>> data, Dataset<String,String> origDataset)Build a Dataset from some data.Dataset<String,String>getDataset(Collection<List<IN>> data)Build a Dataset from some data.Dataset<String,String>getDataset(Collection<List<IN>> data, Index<String> featureIndex, Index<String> classIndex)Build a Dataset from some data.static CMMClassifier<? extends CoreLabel>getDefaultClassifier()Used to obtain the default classifier which is stored inside a jar file.SequenceModelgetSequenceModel(List<IN> document)Set<String>getTags()Returns the Set of entities recognized by this Classifier.voidloadClassifier(ObjectInputStream ois, Properties props)Load a classifier from the given Stream.voidloadDefaultClassifier()Used to load the default supplied classifier.doubleloglikelihood(List<IN> lineInfos)Returns the log conditional likelihood of the given dataset.static voidmain(String[] args)Command-line version of the classifier.Datum<String,String>makeDatum(List<IN> info, int loc, List<FeatureFactory<IN>> featureFactories)Make an individual Datum out of the data list info, focused at position loc.Triple<Counter<Integer>,Counter<Integer>,TwoDimensionalCounter<Integer,String>>printProbsDocument(List<IN> document)voidretrain(ObjectBank<List<IN>> doc)voidretrain(ObjectBank<List<IN>> featureLabels, Index<String> featureIndex, Index<String> labelIndex)Counter<String>scoresOf(List<IN> lineInfos, int pos)voidserializeClassifier(ObjectOutputStream oos)Serialize a sequence classifier to an object output streamvoidserializeClassifier(String serializePath)Serialize a sequence classifier to a file on the given path.voidtrain(Collection<List<IN>> wordInfos, DocumentReaderAndWriter<IN> readerAndWriter)Trains a classifier from a Collection of sequences.voidtrainSemiSup()doubleweight(String feature, String label)double[][]weights()-
Methods inherited from class edu.stanford.nlp.ie.AbstractSequenceClassifier
apply, backgroundSymbol, classify, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswers, classifyAndWriteAnswersKBest, classifyAndWriteAnswersKBest, classifyAndWriteViterbiSearchGraph, classifyFile, classifyFilesAndWriteAnswers, classifyFilesAndWriteAnswers, classifyKBest, classifyRaw, classifySentence, classifySentenceWithGlobalInformation, classifyStdin, classifyStdin, classifyToCharacterOffsets, classifyToString, classifyToString, classifyWithInlineXML, countResults, countResultsSegmenter, defaultReaderAndWriter, dumpFeatures, finalizeClassification, getKnownLCWords, getSampler, labels, loadClassifier, loadClassifier, loadClassifier, loadClassifier, loadClassifier, loadClassifier, loadClassifierNoExceptions, loadClassifierNoExceptions, loadClassifierNoExceptions, loadClassifierNoExceptions, loadClassifierNoExceptions, makeObjectBankFromFile, makeObjectBankFromFile, makeObjectBankFromFiles, makeObjectBankFromFiles, makeObjectBankFromFiles, makeObjectBankFromReader, makeObjectBankFromString, makePlainTextReaderAndWriter, makePlainTextReaderAndWriter, makeReaderAndWriter, plainTextReaderAndWriter, printFeatureLists, printFeatures, printProbs, printProbs, printProbsDocuments, printResults, reinit, segmentString, segmentString, train, train, train, train, train, train, windowSize, writeAnswers
-
-
-
-
Field Detail
-
DEFAULT_CLASSIFIER
public static final String DEFAULT_CLASSIFIER
Default place to look in Jar file for classifier.- See Also:
- Constant Field Values
-
-
Constructor Detail
-
CMMClassifier
protected CMMClassifier()
-
CMMClassifier
public CMMClassifier(Properties props)
-
CMMClassifier
public CMMClassifier(SeqClassifierFlags flags)
-
-
Method Detail
-
getTags
public Set<String> getTags()
Returns the Set of entities recognized by this Classifier.- Returns:
- The Set of entities recognized by this Classifier.
-
classify
public List<IN> classify(List<IN> document)
- Specified by:
classifyin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
document- AListofCoreLabels to be classified.- Returns:
- The same
List, but with the elements annotated with their answers (stored under theCoreAnnotations.AnswerAnnotationkey). The answers will be the class labels defined by the CRF Classifier. They might be things like entity labels (in BIO notation or not) or something like "1" vs. "0" on whether to begin a new token here or not (in word segmentation).
-
classOf
protected String classOf(List<IN> lineInfos, int pos)
Returns the most likely class for the word at the given position.
-
loglikelihood
public double loglikelihood(List<IN> lineInfos)
Returns the log conditional likelihood of the given dataset.- Returns:
- The log conditional likelihood of the given dataset.
-
getSequenceModel
public SequenceModel getSequenceModel(List<IN> document)
- Overrides:
getSequenceModelin classAbstractSequenceClassifier<IN extends CoreLabel>
-
adapt
public void adapt(String filename, Dataset<String,String> trainDataset, DocumentReaderAndWriter<IN> readerWriter)
- Parameters:
filename- adaptation filetrainDataset- original dataset (used in training)
-
adapt
public void adapt(ObjectBank<List<IN>> featureLabels, Dataset<String,String> trainDataset)
- Parameters:
featureLabels- adaptation docstrainDataset- original dataset (used in training)
-
retrain
public void retrain(ObjectBank<List<IN>> featureLabels, Index<String> featureIndex, Index<String> labelIndex)
- Parameters:
featureLabels- retrain docsfeatureIndex- featureIndex of original dataset (used in training)labelIndex- labelIndex of original dataset (used in training)
-
retrain
public void retrain(ObjectBank<List<IN>> doc)
-
train
public void train(Collection<List<IN>> wordInfos, DocumentReaderAndWriter<IN> readerAndWriter)
Description copied from class:AbstractSequenceClassifierTrains a classifier from a Collection of sequences. Note that the Collection can be (and usually is) an ObjectBank.- Specified by:
trainin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
wordInfos- An ObjectBank or a collection of sequences of INreaderAndWriter- A DocumentReaderAndWriter to use when loading test files
-
getDataset
public Dataset<String,String> getDataset(Collection<List<IN>> data)
Build a Dataset from some data. Used for training a classifier.- Parameters:
data- This variable is a list of lists of CoreLabel. That is, it is a collection of documents, each of which is represented as a sequence of CoreLabel objects.- Returns:
- The Dataset which is an efficient encoding of the information in a List of Datums
-
getDataset
public Dataset<String,String> getDataset(Collection<List<IN>> data, Index<String> featureIndex, Index<String> classIndex)
Build a Dataset from some data. Used for training a classifier. By passing in extra featureIndex and classIndex, you can get a Dataset based on featureIndex and classIndex.- Parameters:
data- This variable is a list of lists of CoreLabel. That is, it is a collection of documents, each of which is represented as a sequence of CoreLabel objects.classIndex- if you want to get a Dataset based on featureIndex and classIndex in an existing origDataset- Returns:
- The Dataset which is an efficient encoding of the information in a List of Datums
-
getBiasedDataset
public Dataset<String,String> getBiasedDataset(ObjectBank<List<IN>> data, Index<String> featureIndex, Index<String> classIndex)
-
getDataset
public Dataset<String,String> getDataset(ObjectBank<List<IN>> data, Dataset<String,String> origDataset)
Build a Dataset from some data. Used for training a classifier. By passing in an extra origDataset, you can get a Dataset based on featureIndex and classIndex in an existing origDataset.- Parameters:
data- This variable is a list of lists of CoreLabel. That is, it is a collection of documents, each of which is represented as a sequence of CoreLabel objects.origDataset- if you want to get a Dataset based on featureIndex and classIndex in an existing origDataset- Returns:
- The Dataset which is an efficient encoding of the information in a List of Datums
-
getDataset
public Dataset<String,String> getDataset(Dataset<String,String> oldData, Index<String> goodFeatures)
Build a Dataset from some data.
-
serializeClassifier
public void serializeClassifier(String serializePath)
Description copied from class:AbstractSequenceClassifierSerialize a sequence classifier to a file on the given path.- Specified by:
serializeClassifierin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
serializePath- The path/filename to write the classifier to.
-
serializeClassifier
public void serializeClassifier(ObjectOutputStream oos)
Description copied from class:AbstractSequenceClassifierSerialize a sequence classifier to an object output stream- Specified by:
serializeClassifierin classAbstractSequenceClassifier<IN extends CoreLabel>
-
loadDefaultClassifier
public void loadDefaultClassifier()
Used to load the default supplied classifier. **THIS FUNCTION WILL ONLY WORK IF RUN INSIDE A JAR FILE**
-
getDefaultClassifier
public static CMMClassifier<? extends CoreLabel> getDefaultClassifier()
Used to obtain the default classifier which is stored inside a jar file. THIS FUNCTION WILL ONLY WORK IF RUN INSIDE A JAR FILE.- Returns:
- A Default CMMClassifier from a jar file
-
loadClassifier
public void loadClassifier(ObjectInputStream ois, Properties props) throws ClassCastException, IOException, ClassNotFoundException
Load a classifier from the given Stream. Implementation note: This method does not close the Stream that it reads from.- Specified by:
loadClassifierin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
ois- The ObjectInputStream to load the serialized classifier fromprops- This Properties object will be used to update the SeqClassifierFlags which are read from the serialized classifier- Throws:
IOException- If there are problems accessing the input streamClassCastException- If there are problems interpreting the serialized dataClassNotFoundException- If there are problems interpreting the serialized data
-
getClassifierNoExceptions
public static CMMClassifier<? extends CoreLabel> getClassifierNoExceptions(File file)
-
getClassifier
public static CMMClassifier<? extends CoreLabel> getClassifier(File file) throws IOException, ClassCastException, ClassNotFoundException
-
getClassifierNoExceptions
public static CMMClassifier<CoreLabel> getClassifierNoExceptions(String loadPath)
-
getClassifier
public static CMMClassifier<? extends CoreLabel> getClassifier(String loadPath) throws IOException, ClassCastException, ClassNotFoundException
-
getClassifierNoExceptions
public static CMMClassifier<? extends CoreLabel> getClassifierNoExceptions(InputStream in)
-
getClassifier
public static <INN extends CoreMap> CMMClassifier<? extends CoreLabel> getClassifier(ObjectInputStream ois) throws IOException, ClassCastException, ClassNotFoundException
-
getClassifier
public static <INN extends CoreMap> CMMClassifier<? extends CoreLabel> getClassifier(ObjectInputStream ois, Properties props) throws IOException, ClassCastException, ClassNotFoundException
-
getClassifier
public static CMMClassifier<? extends CoreLabel> getClassifier(InputStream in) throws IOException, ClassCastException, ClassNotFoundException
-
makeDatum
public Datum<String,String> makeDatum(List<IN> info, int loc, List<FeatureFactory<IN>> featureFactories)
Make an individual Datum out of the data list info, focused at position loc.- Parameters:
info- A List of IN objectsloc- The position in the info list to focus feature creation onfeatureFactories- The factory that constructs features out of the item- Returns:
- A Datum (BasicDatum) representing this data instance
-
trainSemiSup
public void trainSemiSup()
-
weights
public double[][] weights()
-
classifyWithGlobalInformation
public List<IN> classifyWithGlobalInformation(List<IN> tokenSeq, CoreMap doc, CoreMap sent)
Description copied from class:AbstractSequenceClassifierClassify aListof something that extendsCoreMapusing as additional information whatever is stored in the document and sentence. This is needed for SUTime (NumberSequenceClassifier), which requires the document date to resolve relative dates.- Specified by:
classifyWithGlobalInformationin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
tokenSeq- AListof something that extendsCoreMap- Returns:
- Classified version of the input tokenSequence
-
printProbsDocument
public Triple<Counter<Integer>,Counter<Integer>,TwoDimensionalCounter<Integer,String>> printProbsDocument(List<IN> document)
Takes aListofCoreLabels and prints the likelihood of each possible label at each point. TODO: Write this method!- Overrides:
printProbsDocumentin classAbstractSequenceClassifier<IN extends CoreLabel>- Parameters:
document- AListofCoreLabels.
-
-