Class ShiftReduceParser

    • Method Detail

      • defaultCoreNLPFlags

        public String[] defaultCoreNLPFlags()
        Description copied from class: ParserGrammar
        Returns a set of options which should be set by default when used in corenlp. For example, the English PCFG/RNN models want -retainTmpSubcategories, and the ShiftReduceParser models may want -beamSize 4 depending on how they were trained.
        TODO: right now completely hardcoded, should be settable as a training time option
        Specified by:
        defaultCoreNLPFlags in class ParserGrammar
      • knownStates

        public Set<String> knownStates()
        Return an unmodifiableSet containing the known states (including binarization)
      • tagSet

        public Set<String> tagSet()
        Return the Set of POS tags used in the model.
      • parse

        public Tree parse​(String sentence)
        Description copied from class: ParserGrammar
        Will parse the text in sentence as if it represented a single sentence by first processing it with a tokenizer.
        Overrides:
        parse in class ParserGrammar
      • parse

        public Tree parse​(List<? extends HasWord> sentence)
        Description copied from class: ParserGrammar
        Parses the list of HasWord. If the parse fails for some reason, an X tree is returned instead of barfing.
        Specified by:
        parse in class ParserGrammar
        Parameters:
        sentence - The input sentence (a List of words)
        Returns:
        A Tree that is the parse tree for the sentence. If the parser fails, a new Tree is synthesized which attaches all words to the root.
      • initialStateFromGoldTagTree

        public static State initialStateFromGoldTagTree​(Tree tree)
      • initialStateFromTaggedSentence

        public static State initialStateFromTaggedSentence​(List<? extends HasWord> words)
      • checkLeafBranching

        public static boolean checkLeafBranching​(Tree tree)
        If an internal node goes directly to a leaf, that is an illegal tree. Otherwise, accept the tree.
        Example:
        (ROOT (sentence (S (morfema.pronominal (PRON Se)) -----(sn (spec (PRON los)) grup.nom)----- (grup.verb (PROPN trago')) (sn (spec (PROPN la)) (grup.nom (DET tierra)))) (NOUN ....) (S (sadv (grup.adv (PROPN ya))) (neg (ADV no)) (grup.verb (ADV viven)) (sp (prep (VERB en)) (sn (grup.nom (ADP Cuba)))) (PROPN ...)) (conj (PUNCT y)) (S (sp (prep (CCONJ a)) (sn (grup.nom (ADP nadie)))) (sn (grup.nom (PRON le))) (grup.verb (PRON importa)) (sn (spec (VERB la)) (grup.nom (S (relatiu (DET que)) (grup.verb (PRON esten) (gerundi (VERB pasando))))))) (VERB ....)))
      • checkRootTransition

        public static boolean checkRootTransition​(Tree tree)
        Some trees in the English datasets have a binary transition at the top, which we don't like as it teaches the parser to sometimes make binary transitions in normal trees
      • filterTreebank

        public List<Tree> filterTreebank​(Treebank treebank)
        Filters any trees which are obviously unacceptable for the sr parser training.
        Disallowed:
        • trees where internal nodes go directly to leaves instead of preterminals.
        • trees which don't start with a unary transition
      • findKnownStates

        public static Set<String> findKnownStates​(List<Tree> binarizedTrees)
      • findKnownStates

        public static void findKnownStates​(Tree tree,
                                           Set<String> knownStates)
      • redoTags

        public static void redoTags​(Tree tree,
                                    Tagger tagger)
      • redoTags

        public static void redoTags​(List<Tree> trees,
                                    Tagger tagger,
                                    int nThreads)
      • saveModel

        public void saveModel​(String path)
      • main

        public static void main​(String[] args)