Class ShiftReduceParser
- java.lang.Object
-
- edu.stanford.nlp.parser.common.ParserGrammar
-
- edu.stanford.nlp.parser.shiftreduce.ShiftReduceParser
-
- All Implemented Interfaces:
ParserQueryFactory,Serializable,Function<List<? extends HasWord>,Tree>
public class ShiftReduceParser extends ParserGrammar implements Serializable
A shift-reduce constituency parser. Overview and description available at https://nlp.stanford.edu/software/srparser.shtml- Author:
- John Bauer
- See Also:
- Serialized Form
-
-
Constructor Summary
Constructors Constructor Description ShiftReduceParser(ShiftReduceOptions op)ShiftReduceParser(ShiftReduceOptions op, PerceptronModel model)
-
Method Summary
All Methods Static Methods Instance Methods Concrete Methods Modifier and Type Method Description static List<Tree>binarizeTreebank(Iterable<Tree> treebank, Options op)static ShiftReduceOptionsbuildTrainingOptions(String tlppClass, String[] args)static booleancheckLeafBranching(Tree tree)If an internal node goes directly to a leaf, that is an illegal tree.static booleancheckRootTransition(Tree tree)Some trees in the English datasets have a binary transition at the top, which we don't like as it teaches the parser to sometimes make binary transitions in normal treesString[]defaultCoreNLPFlags()Returns a set of options which should be set by default when used in corenlp.List<Tree>filterTreebank(Treebank treebank)Filters any trees which are obviously unacceptable for the sr parser training.static voidfindKnownStates(Tree tree, Set<String> knownStates)static Set<String>findKnownStates(List<Tree> binarizedTrees)List<Eval>getExtraEvals()TODO: add an eval which measures transition accuracy?OptionsgetOp()List<ParserQueryEval>getParserQueryEvals()Return a list of Eval-style objects which care about the whole ParserQuery, not just the finished treeTreebankLangParserParamsgetTLPParams()static StateinitialStateFromGoldTagTree(Tree tree)static StateinitialStateFromTaggedSentence(List<? extends HasWord> words)Set<String>knownStates()Return an unmodifiableSet containing the known states (including binarization)static ShiftReduceParserloadModel(String path, String... extraFlags)static voidmain(String[] args)Treeparse(String sentence)Will parse the text insentenceas if it represented a single sentence by first processing it with a tokenizer.Treeparse(List<? extends HasWord> sentence)Parses the list of HasWord.ParserQueryparserQuery()TreeparseTree(List<? extends HasWord> sentence)Similar to parse(), but instead of returning an X tree on failure, returns null.List<Tree>readBinarizedTreebank(String treebankPath, FileFilter treebankFilter)TreebankreadTreebank(String treebankPath, FileFilter treebankFilter)static voidredoTags(Tree tree, Tagger tagger)static voidredoTags(List<Tree> trees, Tagger tagger, int nThreads)booleanrequiresTags()The model requires text to be pretaggedvoidsaveModel(String path)voidsetOptionFlags(String... flags)Set<String>tagSet()Return the Set of POS tags used in the model.TreebankLanguagePacktreebankLanguagePack()-
Methods inherited from class edu.stanford.nlp.parser.common.ParserGrammar
apply, lemmatize, lemmatize, loadModelFromZip, loadTagger, tokenize
-
-
-
-
Constructor Detail
-
ShiftReduceParser
public ShiftReduceParser(ShiftReduceOptions op)
-
ShiftReduceParser
public ShiftReduceParser(ShiftReduceOptions op, PerceptronModel model)
-
-
Method Detail
-
getOp
public Options getOp()
- Specified by:
getOpin classParserGrammar
-
getTLPParams
public TreebankLangParserParams getTLPParams()
- Specified by:
getTLPParamsin classParserGrammar
-
treebankLanguagePack
public TreebankLanguagePack treebankLanguagePack()
- Specified by:
treebankLanguagePackin classParserGrammar
-
defaultCoreNLPFlags
public String[] defaultCoreNLPFlags()
Description copied from class:ParserGrammarReturns a set of options which should be set by default when used in corenlp. For example, the English PCFG/RNN models want -retainTmpSubcategories, and the ShiftReduceParser models may want -beamSize 4 depending on how they were trained.
TODO: right now completely hardcoded, should be settable as a training time option- Specified by:
defaultCoreNLPFlagsin classParserGrammar
-
knownStates
public Set<String> knownStates()
Return an unmodifiableSet containing the known states (including binarization)
-
requiresTags
public boolean requiresTags()
Description copied from class:ParserGrammarThe model requires text to be pretagged- Specified by:
requiresTagsin classParserGrammar
-
parserQuery
public ParserQuery parserQuery()
- Specified by:
parserQueryin interfaceParserQueryFactory- Specified by:
parserQueryin classParserGrammar
-
parse
public Tree parse(String sentence)
Description copied from class:ParserGrammarWill parse the text insentenceas if it represented a single sentence by first processing it with a tokenizer.- Overrides:
parsein classParserGrammar
-
parse
public Tree parse(List<? extends HasWord> sentence)
Description copied from class:ParserGrammarParses the list of HasWord. If the parse fails for some reason, an X tree is returned instead of barfing.- Specified by:
parsein classParserGrammar- Parameters:
sentence- The input sentence (a List of words)- Returns:
- A Tree that is the parse tree for the sentence. If the parser fails, a new Tree is synthesized which attaches all words to the root.
-
parseTree
public Tree parseTree(List<? extends HasWord> sentence)
Description copied from class:ParserGrammarSimilar to parse(), but instead of returning an X tree on failure, returns null.- Specified by:
parseTreein classParserGrammar
-
getExtraEvals
public List<Eval> getExtraEvals()
TODO: add an eval which measures transition accuracy?- Specified by:
getExtraEvalsin classParserGrammar
-
getParserQueryEvals
public List<ParserQueryEval> getParserQueryEvals()
Description copied from class:ParserGrammarReturn a list of Eval-style objects which care about the whole ParserQuery, not just the finished tree- Specified by:
getParserQueryEvalsin classParserGrammar
-
initialStateFromTaggedSentence
public static State initialStateFromTaggedSentence(List<? extends HasWord> words)
-
buildTrainingOptions
public static ShiftReduceOptions buildTrainingOptions(String tlppClass, String[] args)
-
readTreebank
public Treebank readTreebank(String treebankPath, FileFilter treebankFilter)
-
readBinarizedTreebank
public List<Tree> readBinarizedTreebank(String treebankPath, FileFilter treebankFilter)
-
checkLeafBranching
public static boolean checkLeafBranching(Tree tree)
If an internal node goes directly to a leaf, that is an illegal tree. Otherwise, accept the tree.
Example:(ROOT (sentence (S (morfema.pronominal (PRON Se)) -----(sn (spec (PRON los)) grup.nom)----- (grup.verb (PROPN trago')) (sn (spec (PROPN la)) (grup.nom (DET tierra)))) (NOUN ....) (S (sadv (grup.adv (PROPN ya))) (neg (ADV no)) (grup.verb (ADV viven)) (sp (prep (VERB en)) (sn (grup.nom (ADP Cuba)))) (PROPN ...)) (conj (PUNCT y)) (S (sp (prep (CCONJ a)) (sn (grup.nom (ADP nadie)))) (sn (grup.nom (PRON le))) (grup.verb (PRON importa)) (sn (spec (VERB la)) (grup.nom (S (relatiu (DET que)) (grup.verb (PRON esten) (gerundi (VERB pasando))))))) (VERB ....)))
-
checkRootTransition
public static boolean checkRootTransition(Tree tree)
Some trees in the English datasets have a binary transition at the top, which we don't like as it teaches the parser to sometimes make binary transitions in normal trees
-
filterTreebank
public List<Tree> filterTreebank(Treebank treebank)
Filters any trees which are obviously unacceptable for the sr parser training.
Disallowed:- trees where internal nodes go directly to leaves instead of preterminals.
- trees which don't start with a unary transition
-
setOptionFlags
public void setOptionFlags(String... flags)
- Specified by:
setOptionFlagsin classParserGrammar
-
loadModel
public static ShiftReduceParser loadModel(String path, String... extraFlags)
-
saveModel
public void saveModel(String path)
-
main
public static void main(String[] args)
-
-