Challenge: Discontinuous constituency parsing is still being developed for its efficiency and accuracy are far behind its continuous counterparts.
Approach: They propose to transform a discontinuous constituent tree into a pseudo-continuous one by reordering words in the sentence.
Outcome: The proposed method can transform a discontinuous constituent tree into a pseudo-continuous one by parsing and performing actions on three classical discontinuous constituency treebanks.

Similar Papers

Discontinuous Combinatory Constituency Parsing (2023.tacl-1)

Copied to clipboard

Challenge: Discontinuous parsing is more challenging than continuous parsers because children can group with syntactic cousins in the sentence rather than its two adjacent neighbors.
Approach: They extend a pair of combinator-based constituency parsers into a discontinuous pair . they use a swap action and biaffine attention to iteratively compose constituent vectors from word embeddings without any grammar constraints.
Outcome: The proposed parsers achieve state-of-the-art discontinuous accuracy with a significant speed advantage over continuous parsing.
Reducing Discontinuous to Continuous Parsing with Pointer Network Reordering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing discontinuous constituent parsers are slow and lack accuracy and speed . however, discontinuous parsing can be solved by any off-the-shelf continuous parser .
Approach: They propose to reduce discontinuous constituent parsing to a continuous problem by reordering tokens.
Outcome: The proposed method is on par with state-of-the-art methods but considerably faster.
Discontinuous Constituency Parsing with a Stack-Free Transition System and a Dynamic Oracle (N19-1)

Copied to clipboard

Challenge: Discontinuous constituency trees are derivations of Linear Context-Free Rewriting Systems (LCFRS), which makes them much harder to parse.
Approach: They propose a transition system that uses a set of parsing items with constant-time random access instead of storing subtrees in a stack .
Outcome: The proposed system constructs a discontinuous constituency tree in 4n–2 transitions for a sentence of length n.
Don’t Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span Selection (2023.acl-long)

Copied to clipboard

Challenge: Constituency parsing is a fundamental task in natural language processing, having many applications in downstream tasks such as language modeling.
Approach: They propose a simple and unified approach for both continuous and discontinuous constituency parsing via autoregressive span selection.
Outcome: The proposed model can predict all possible continuous and discontinuous constituency trees without sacrificing data coverage and without expensive chart-based parsing algorithms.
Span-based discontinuous constituency parsing: a family of exact chart-based algorithms with time complexities from O(nˆ6) down to O(nˆ3) (2020.emnlp-main)

Copied to clipboard

Challenge: a novel chart-based parser for discontinuous constituency trees is proposed for span-based span parsing . it can process discontinuous constituent trees of block degree two, including ill-nested structures .
Approach: They propose a chart-based algorithm for span-based parsing of discontinuous constituency trees . they build variants with smaller search spaces and time complexities ranging from O(n6) down to O(N3) .
Outcome: The proposed algorithm can process 98% of constituents in linguistic treebanks while having the same complexity as continuous constituency parsers.
BERT-Proof Syntactic Structures: Investigating Errors in Discontinuous Constituency Parsing (2021.findings-acl)

Copied to clipboard

Challenge: Recent results show that pretrained language models can be used for many tasks with high accuracy and high performance.
Approach: They propose two methods for automatically analysing discontinuous parsers' errors.
Outcome: The proposed methods characterize errors of a state-of-the-art transition-based discontinuous parser and provide an overview of the contribution of BERT to this task.
Discontinuous Constituent Parsing as Sequence Labeling (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to discontinuous parsing are complex and low-level.
Approach: They propose to encode discontinuities as nearly ordered permutations of the input sequence.
Outcome: The proposed model is fast and accurate under the right representation.
Straight to the Tree: Constituency Parsing with Neural Syntactic Distance (P18-1)

Copied to clipboard

Challenge: Compared to traditional shift-reduce parsing schemes, our approach is free from the potentially disastrous compounding error.
Approach: They propose a model that predicts a scalar for each split position in a sentence and then determines the topology of grammar tree based on syntactic distances.
Outcome: The proposed model achieves the state-of-the-art single model F1 score of 92.1 on PTB and 86.4 on CTB dataset, surpassing the previous single model results by a large margin.
An Empirical Study of Building a Strong Baseline for Constituency Parsing (P18-2)

Copied to clipboard

Challenge: Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization.
Approach: They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing.
Outcome: The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing.
Parsing Headed Constituencies (2024.lrec-main)

Copied to clipboard

Challenge: Using constituency and dependency trees, syntactic representations are preferred for tasks such as nominal phrase extraction and identification of terminology.
Approach: They propose a parsing technique that generates headed constituency trees which combine information typically contained in constituency and dependency trees.
Outcome: The proposed method generates headed constituency trees with discontinuities and can generate constituency tree with discontinuity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations