Class ChineseMaxentLexicon
- java.lang.Object
-
- edu.stanford.nlp.parser.lexparser.ChineseMaxentLexicon
-
- All Implemented Interfaces:
Lexicon,Serializable
public class ChineseMaxentLexicon extends Object implements Lexicon
A Lexicon class that computes the score of word|tag according to a maxent model of tag|word (divided by MLE estimate of P(tag)).
It would be nice to factor out a superclass MaxentLexicon that takes a WordFeatureExtractor- Author:
- Galen Andrew
- See Also:
- Serialized Form
-
-
Field Summary
Fields Modifier and Type Field Description static booleanfixUnkFunctionWordsstatic booleanseenTagsOnlyCollectionValuedMap<String,String>tagsForWord-
Fields inherited from interface edu.stanford.nlp.parser.lexparser.Lexicon
BOUNDARY, BOUNDARY_TAG, UNKNOWN_WORD
-
-
Method Summary
All Methods Static Methods Instance Methods Concrete Methods Modifier and Type Method Description voidfinishTraining()Done collecting statistics for the lexicon.UnknownWordModelgetUnknownWordModel()voidincrementTreesRead(double weight)If training on a per-word basis instead of on a per-tree basis, we will want to increment the tree count as this happens.voidinitializeTraining(double numTrees)Start training this lexicon on the expected number of trees.booleanisKnown(int word)Checks whether a word is in the lexicon.booleanisKnown(String word)Checks whether a word is in the lexicon.static voidmain(String[] args)intnumRules()Returns the number of rules (tag rewrites as word) in the Lexicon.voidreadData(BufferedReader in)Read the lexicon from the BufferedReader in the format written by writeData.Iterator<IntTaggedWord>ruleIteratorByWord(int word, int loc, String featureSpec)Get an iterator over all rules (pairs of (word, POS)) for this word.Iterator<IntTaggedWord>ruleIteratorByWord(String word, int loc, String featureSpec)Same thing, but with a string that needs to be translated by the lexicon's word indexfloatscore(IntTaggedWord iTW, int loc, String word, String featureSpec)Get the score of this word with this tag (as an IntTaggedWord) at this loc.voidsetUnknownWordModel(UnknownWordModel uwm)Set<String>tagSet(Function<String,String> basicCategoryFunction)Return the Set of tags used by this tagger (available after training the tagger).voidtrain(TaggedWord tw, int loc, double weight)Not all subclasses support this particular method.voidtrain(Tree tree, double weight)Add the given tree to the statistics counted.voidtrain(Collection<Tree> trees)Add the given collection of trees to the statistics counted.voidtrain(Collection<Tree> trees, double weight)Add the given collection of trees to the statistics counted.voidtrain(Collection<Tree> trees, Collection<Tree> rawTrees)voidtrain(List<TaggedWord> sentence, double weight)Add the given sentence to the statistics counted.voidtrainUnannotated(List<TaggedWord> sentence, double weight)Sometimes we might have a sentence of tagged words which we would like to add to the lexicon, but they weren't part of a binarized, markovized, or otherwise annotated tree.voidwriteData(Writer w)Write the lexicon in human-readable format to the Writer.
-
-
-
Field Detail
-
seenTagsOnly
public static final boolean seenTagsOnly
- See Also:
- Constant Field Values
-
fixUnkFunctionWords
public static final boolean fixUnkFunctionWords
- See Also:
- Constant Field Values
-
tagsForWord
public CollectionValuedMap<String,String> tagsForWord
-
-
Method Detail
-
isKnown
public boolean isKnown(int word)
Description copied from interface:LexiconChecks whether a word is in the lexicon.
-
isKnown
public boolean isKnown(String word)
Description copied from interface:LexiconChecks whether a word is in the lexicon.
-
tagSet
public Set<String> tagSet(Function<String,String> basicCategoryFunction)
Return the Set of tags used by this tagger (available after training the tagger).
-
ruleIteratorByWord
public Iterator<IntTaggedWord> ruleIteratorByWord(int word, int loc, String featureSpec)
Description copied from interface:LexiconGet an iterator over all rules (pairs of (word, POS)) for this word.- Specified by:
ruleIteratorByWordin interfaceLexicon- Parameters:
word- The word, represented as an integer in Indexloc- The position of the word in the sentence (counting from 0). Implementation note: The BaseLexicon class doesn't actually make use of this position information.featureSpec- Additional word features like morphosyntactic information.- Returns:
- An Iterator over a List ofIntTaggedWords, which pair the word
with possible taggings as integer pairs. (Each can be
thought of as a
tag -> wordrule.)
-
ruleIteratorByWord
public Iterator<IntTaggedWord> ruleIteratorByWord(String word, int loc, String featureSpec)
Description copied from interface:LexiconSame thing, but with a string that needs to be translated by the lexicon's word index- Specified by:
ruleIteratorByWordin interfaceLexicon
-
numRules
public int numRules()
Returns the number of rules (tag rewrites as word) in the Lexicon. This method isn't yet implemented in this class. It currently just returns 0, which may or may not be helpful.
-
initializeTraining
public void initializeTraining(double numTrees)
Description copied from interface:LexiconStart training this lexicon on the expected number of trees. (Some UnknownWordModels use the number of trees to know when to start counting statistics.)- Specified by:
initializeTrainingin interfaceLexicon
-
train
public final void train(Collection<Tree> trees)
Add the given collection of trees to the statistics counted. Can be called multiple times with different trees.
-
train
public void train(Collection<Tree> trees, double weight)
Add the given collection of trees to the statistics counted. Can be called multiple times with different trees.
-
train
public void train(Tree tree, double weight)
Add the given tree to the statistics counted. Can be called multiple times with different trees.
-
train
public void train(List<TaggedWord> sentence, double weight)
Add the given sentence to the statistics counted. Can be called multiple times with different sentences.
-
trainUnannotated
public void trainUnannotated(List<TaggedWord> sentence, double weight)
Description copied from interface:LexiconSometimes we might have a sentence of tagged words which we would like to add to the lexicon, but they weren't part of a binarized, markovized, or otherwise annotated tree.- Specified by:
trainUnannotatedin interfaceLexicon
-
incrementTreesRead
public void incrementTreesRead(double weight)
Description copied from interface:LexiconIf training on a per-word basis instead of on a per-tree basis, we will want to increment the tree count as this happens.- Specified by:
incrementTreesReadin interfaceLexicon
-
train
public void train(TaggedWord tw, int loc, double weight)
Description copied from interface:LexiconNot all subclasses support this particular method. Those that don't will barf...
-
finishTraining
public void finishTraining()
Description copied from interface:LexiconDone collecting statistics for the lexicon.- Specified by:
finishTrainingin interfaceLexicon
-
main
public static void main(String[] args)
-
score
public float score(IntTaggedWord iTW, int loc, String word, String featureSpec)
Description copied from interface:LexiconGet the score of this word with this tag (as an IntTaggedWord) at this loc. (Presumably an estimate of P(word | tag).)- Specified by:
scorein interfaceLexicon- Parameters:
iTW- An IntTaggedWord pairing a word and POS tagloc- The position in the sentence. In the default implementation this is used only for unknown words to change their probability distribution when sentence initial.word- The word itself; useful so we don't have to look it up in an indexfeatureSpec- TODO- Returns:
- A score, usually, log P(word|tag)
-
writeData
public void writeData(Writer w) throws IOException
Description copied from interface:LexiconWrite the lexicon in human-readable format to the Writer. (An optional operation.)- Specified by:
writeDatain interfaceLexicon- Parameters:
w- The writer to output to- Throws:
IOException- If any I/O problem
-
readData
public void readData(BufferedReader in) throws IOException
Description copied from interface:LexiconRead the lexicon from the BufferedReader in the format written by writeData. (An optional operation.)- Specified by:
readDatain interfaceLexicon- Parameters:
in- The BufferedReader to read from- Throws:
IOException- If any I/O problem
-
getUnknownWordModel
public UnknownWordModel getUnknownWordModel()
- Specified by:
getUnknownWordModelin interfaceLexicon
-
setUnknownWordModel
public void setUnknownWordModel(UnknownWordModel uwm)
- Specified by:
setUnknownWordModelin interfaceLexicon
-
train
public void train(Collection<Tree> trees, Collection<Tree> rawTrees)
-
-