Class ChineseMaxentLexicon

  • All Implemented Interfaces:
    Lexicon, Serializable

    public class ChineseMaxentLexicon
    extends Object
    implements Lexicon
    A Lexicon class that computes the score of word|tag according to a maxent model of tag|word (divided by MLE estimate of P(tag)).
    It would be nice to factor out a superclass MaxentLexicon that takes a WordFeatureExtractor
    Author:
    Galen Andrew
    See Also:
    Serialized Form
    • Constructor Detail

    • Method Detail

      • isKnown

        public boolean isKnown​(int word)
        Description copied from interface: Lexicon
        Checks whether a word is in the lexicon.
        Specified by:
        isKnown in interface Lexicon
        Parameters:
        word - The word as an int
        Returns:
        Whether the word is in the lexicon
      • isKnown

        public boolean isKnown​(String word)
        Description copied from interface: Lexicon
        Checks whether a word is in the lexicon.
        Specified by:
        isKnown in interface Lexicon
        Parameters:
        word - The word as a String
        Returns:
        Whether the word is in the lexicon
      • tagSet

        public Set<String> tagSet​(Function<String,​String> basicCategoryFunction)
        Return the Set of tags used by this tagger (available after training the tagger).
        Specified by:
        tagSet in interface Lexicon
        Returns:
        The Set of tags used by this tagger
      • ruleIteratorByWord

        public Iterator<IntTaggedWord> ruleIteratorByWord​(int word,
                                                          int loc,
                                                          String featureSpec)
        Description copied from interface: Lexicon
        Get an iterator over all rules (pairs of (word, POS)) for this word.
        Specified by:
        ruleIteratorByWord in interface Lexicon
        Parameters:
        word - The word, represented as an integer in Index
        loc - The position of the word in the sentence (counting from 0). Implementation note: The BaseLexicon class doesn't actually make use of this position information.
        featureSpec - Additional word features like morphosyntactic information.
        Returns:
        An Iterator over a List ofIntTaggedWords, which pair the word with possible taggings as integer pairs. (Each can be thought of as a tag -> word rule.)
      • numRules

        public int numRules()
        Returns the number of rules (tag rewrites as word) in the Lexicon. This method isn't yet implemented in this class. It currently just returns 0, which may or may not be helpful.
        Specified by:
        numRules in interface Lexicon
        Returns:
        The number of rules (tag rewrites as word) in the Lexicon.
      • initializeTraining

        public void initializeTraining​(double numTrees)
        Description copied from interface: Lexicon
        Start training this lexicon on the expected number of trees. (Some UnknownWordModels use the number of trees to know when to start counting statistics.)
        Specified by:
        initializeTraining in interface Lexicon
      • train

        public final void train​(Collection<Tree> trees)
        Add the given collection of trees to the statistics counted. Can be called multiple times with different trees.
        Specified by:
        train in interface Lexicon
        Parameters:
        trees - Trees to train on
      • train

        public void train​(Collection<Tree> trees,
                          double weight)
        Add the given collection of trees to the statistics counted. Can be called multiple times with different trees.
        Specified by:
        train in interface Lexicon
      • train

        public void train​(Tree tree,
                          double weight)
        Add the given tree to the statistics counted. Can be called multiple times with different trees.
        Specified by:
        train in interface Lexicon
      • train

        public void train​(List<TaggedWord> sentence,
                          double weight)
        Add the given sentence to the statistics counted. Can be called multiple times with different sentences.
        Specified by:
        train in interface Lexicon
      • trainUnannotated

        public void trainUnannotated​(List<TaggedWord> sentence,
                                     double weight)
        Description copied from interface: Lexicon
        Sometimes we might have a sentence of tagged words which we would like to add to the lexicon, but they weren't part of a binarized, markovized, or otherwise annotated tree.
        Specified by:
        trainUnannotated in interface Lexicon
      • incrementTreesRead

        public void incrementTreesRead​(double weight)
        Description copied from interface: Lexicon
        If training on a per-word basis instead of on a per-tree basis, we will want to increment the tree count as this happens.
        Specified by:
        incrementTreesRead in interface Lexicon
      • train

        public void train​(TaggedWord tw,
                          int loc,
                          double weight)
        Description copied from interface: Lexicon
        Not all subclasses support this particular method. Those that don't will barf...
        Specified by:
        train in interface Lexicon
      • finishTraining

        public void finishTraining()
        Description copied from interface: Lexicon
        Done collecting statistics for the lexicon.
        Specified by:
        finishTraining in interface Lexicon
      • main

        public static void main​(String[] args)
      • score

        public float score​(IntTaggedWord iTW,
                           int loc,
                           String word,
                           String featureSpec)
        Description copied from interface: Lexicon
        Get the score of this word with this tag (as an IntTaggedWord) at this loc. (Presumably an estimate of P(word | tag).)
        Specified by:
        score in interface Lexicon
        Parameters:
        iTW - An IntTaggedWord pairing a word and POS tag
        loc - The position in the sentence. In the default implementation this is used only for unknown words to change their probability distribution when sentence initial.
        word - The word itself; useful so we don't have to look it up in an index
        featureSpec - TODO
        Returns:
        A score, usually, log P(word|tag)
      • writeData

        public void writeData​(Writer w)
                       throws IOException
        Description copied from interface: Lexicon
        Write the lexicon in human-readable format to the Writer. (An optional operation.)
        Specified by:
        writeData in interface Lexicon
        Parameters:
        w - The writer to output to
        Throws:
        IOException - If any I/O problem
      • readData

        public void readData​(BufferedReader in)
                      throws IOException
        Description copied from interface: Lexicon
        Read the lexicon from the BufferedReader in the format written by writeData. (An optional operation.)
        Specified by:
        readData in interface Lexicon
        Parameters:
        in - The BufferedReader to read from
        Throws:
        IOException - If any I/O problem