Papers with PDTB
Towards Unification of Discourse Annotation Frameworks (2022.acl-srw)
Copied to clipboard
| Challenge: | Discourse information is difficult to represent and annotate, and corpora annotated under different frameworks vary considerably. |
| Approach: | They propose to use automatic means to unify discourse structures and relations . they will also explore the application of the unified framework in multi-task learning and graphical models . |
| Outcome: | The proposed method can be used in multi-task learning and graphical models. |
Improving Implicit Discourse Relation Classification by Modeling Inter-dependencies of Discourse Units in a Paragraph (N18-1)
Copied to clipboard
| Challenge: | Existing methods for predicting implicit discourse relations ignore wider paragraph contexts beyond the two discourse units examined for a discourse relation prediction. |
| Approach: | They propose a paragraph-level neural network that models inter-dependencies between discourse units and discourse relation continuity and patterns and predicts a sequence of discourse relations in a sentence. |
| Outcome: | The proposed model outperforms state-of-the-art systems on the benchmark corpus of PDTB. |
TransS-Driven Joint Learning Architecture for Implicit Discourse Relation Recognition (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to implicit discourse relation recognition lack connectives as strong linguistic clues. |
| Approach: | They propose a transS-driven joint learning architecture to translate discourse relations in low-dimensional embedding space and exploit the semantic features of arguments to assist discourse understanding. |
| Outcome: | The proposed model outperforms existing systems on the Penn Discourse TreeBank. |
Implicit Discourse Relation Recognition using Neural Tensor Network with Interactive Attention and Sparse Learning (C18-1)
Copied to clipboard
| Challenge: | Existing methods for implicit discourse relation recognition ignore bidirectional interactions between two arguments and sparsity of pair patterns. |
| Approach: | They propose a neural Tensor network framework with interactive attention and sparse learning for implicit discourse relation recognition. |
| Outcome: | The proposed framework is effective on PDTB and can be used in text summarization, conversation system and so on. |
Towards Identifying Alternative-Lexicalization Signals of Discourse Relations (2022.coling-1)
Copied to clipboard
| Challenge: | Existing shallow discourse parsing methods have been limited to identifying relations signaled by a discourse connective and those without a signal. |
| Approach: | They propose to identify relations signalled by a discourse connective and those without . they compare a pattern-based approach and a sequence labeling model . |
| Outcome: | The proposed approach is based on a pattern-based approach and a sequence labeling model. |
SciDTB: Discourse Dependency TreeBank for Scientific Abstracts (P18-2)
Copied to clipboard
| Challenge: | Discourse relations are annotated on scientific articles. |
| Approach: | They propose a domain-specific discourse treebank annotated on scientific articles . they use dependency trees to represent discourse structure, which is flexible and simplified . |
| Outcome: | The proposed treebank is a benchmark for evaluating discourse dependency parsers. |
A Survey of QUD Models for Discourse Processing (2025.naacl-long)
Copied to clipboard
| Challenge: | Question Under Discussion (QUD) is a linguistic analytic framework for explaining pragmatic phenomena and information structural analysis. |
| Approach: | They propose to use Question Under Discussion (QUD) to model discourse units, such as sentences, as answers to some implicit or explicit questions. |
| Outcome: | The proposed model is compared with RST, PDTB and SDRT . questions that may require further study are suggested. |
Entity Enhancement for Implicit Discourse Relation Classification in the Biomedical Domain (2021.acl-short)
Copied to clipboard
| Challenge: | Discourse relation classification is a challenging task when the text domain is different from the standard Penn Discourse Treebank (PDTB) training corpus domain. |
| Approach: | They propose to use the Biomedical Discourse Relation Bank to improve discourse relational argument representation by linking explicit instances of similar relations with a voting pipeline. |
| Outcome: | The proposed model outperforms the pre-trained BioBERT model by 2% points. |
Shallow Discourse Parsing for Under-Resourced Languages: Combining Machine Translation and Annotation Projection (2020.lrec-1)
Copied to clipboard
| Challenge: | Shallow Discourse Parsing (SDP) relies on large amounts of training data, which so far exists only for English. |
| Approach: | They propose to translate an existing English Penn Discourse TreeBank into German and use it to create a German corpus annotated for shallow discourse relations in the news domain. |
| Outcome: | The proposed corpus is annotated for shallow discourse relations in the (financial) news domain. |
Using a Penalty-based Loss Re-estimation Method to Improve Implicit Discourse Relation Classification (2020.coling-main)
Copied to clipboard
| Challenge: | inessential words are unintentionally misjudged as attention-worthy words and assigned heavier attention weights than should be. |
| Approach: | They propose a penalty-based method to regulate the attention learning process by integrating penalty coefficients into the computation of loss by means of overstability of attention weight distributions. |
| Outcome: | The proposed method improves on the Penn Discourse TreeBank corpus and is competitive compared to the state-of-the-art methods. |
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)
Copied to clipboard
| Challenge: | Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels. |
| Approach: | They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus. |
| Outcome: | The proposed corpus is more usable for shallow discourse parsing. |
What Causes the Failure of Explicit to Implicit Discourse Relation Recognition? (2024.naacl-long)
Copied to clipboard
| Challenge: | Prior work claimed that explicit classifiers perform poorly in implicit scenarios . a label shift occurs after connectives are removed, but no empirical evidence supports this claim . |
| Approach: | They propose to prove that the discourse relations expressed by some explicit instances will change when connectives disappear. |
| Outcome: | The proposed methods outperform strong baselines on PDTB 2.0, PDTT 3.0, and the GUM dataset. |
The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)
Copied to clipboard
Fiona Anting Tan, Ali Hürriyetoğlu, Tommaso Caselli, Nelleke Oostdijk, Tadashi Nomoto, Hansi Hettiarachchi, Iqra Ameer, Onur Uca, Farhana Ferdousi Liza, Tiancheng Hu
| Challenge: | Existing annotation guidelines for event causality focus on only explicit relations or clauses. |
| Approach: | They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not. |
| Outcome: | The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation . |
Inducing Discourse Marker Inventories from Lexical Knowledge Graphs (2022.lrec-1)
Copied to clipboard
| Challenge: | Discourse marker inventories are important tools for the development of discourse parsers and corpora with discourse annotations. |
| Approach: | They explore the potential of multilingual lexical knowledge graphs to induce multilingual discourse marker lexicons using concept propagation methods previously developed in translation inference across dictionaries. |
| Outcome: | The proposed method can induce multilingual discourse marker lexicons using multilingual knowledge graphs. |
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)
Copied to clipboard
| Challenge: | Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications. |
| Approach: | They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages. |
| Outcome: | The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages. |
Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models’ Understanding of Discourse Relations (2024.acl-long)
Copied to clipboard
| Challenge: | Discursive Socratic Questioning (DISQ) assesses a model's understanding of discourse relations by requiring systematic accuracy over multiple questions. |
| Approach: | They propose a method that evaluates faithfulness of understanding discourse based on question answering. |
| Outcome: | The proposed method evaluates the faithfulness of understanding discourse based on question answering. |
Enhancing Reasoning Capabilities by Instruction Learning and Chain-of-Thoughts for Implicit Discourse Relation Recognition (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for implicit discourse relation recognition are based on generative models, but some studies suggest they do not perform as well as generic encoder-only models for NLU tasks. |
| Approach: | They propose a classification method that is solely based on generative models and utilize Chain-of-Thoughts to partition the inference process into a sequence of three successive stages. |
| Outcome: | The proposed model outperforms existing models on a natural language understanding task. |
DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing (2024.lrec-main)
Copied to clipboard
Chloé Braud, Amir Zeldes, Laura Rivière, Yang Janet Liu, Philippe Muller, Damien Sileo, Tatsuya Aoyama
| Challenge: | DISRPT is a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing. |
| Approach: | They present a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing that includes 13 languages and 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks. |
| Outcome: | The DISRPT dataset includes data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks. |
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature. |
| Approach: | They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders . |
| Outcome: | The proposed evaluation protocol improves the existing framework and provides strong baseline results. |
Multi-Label Classification for Implicit Discourse Relation Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Prior research in discourse relation recognition has treated these instances as separate examples during training, with a gold-standard prediction matching one of the labels considered correct at test time. |
| Approach: | They propose to use multiple labels to annotate an example when multiple relations are believed to hold simultaneously. |
| Outcome: | The proposed frameworks don't depress performance for single-label prediction. |
Global and Local Hierarchy-aware Contrastive Framework for Implicit Discourse Relation Recognition (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to integrate whole hierarchical information of senses into discourse relation representations for multi-level sense recognition ignore static hierarchic structure containing all senses and ignore hierarchically sense label sequence corresponding to each instance. |
| Approach: | They propose to use a GlObal and Local Hierarchy-aware Contrastive Framework to model two kinds of hierarchies with the aid of multi-task learning and contrastive learning to learn better representations of discourse relation relationships. |
| Outcome: | The proposed method outperforms current state-of-the-art models at all hierarchical levels on PDTB 2.0 and PDTP 3.0 datasets. |
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)
Copied to clipboard
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Levine, Jessica Lin, Devika Tiwari, Amir Zeldes
| Challenge: | Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old. |
| Approach: | They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing. |
| Outcome: | The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain. |
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)
Copied to clipboard
| Challenge: | Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey. |
| Approach: | They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text. |
| Outcome: | The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications. |
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)
Copied to clipboard
| Challenge: | lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels. |
| Approach: | They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense. |
| Outcome: | The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format. |
Annotation-Inspired Implicit Discourse Relation Classification with Auxiliary Discourse Connective Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Discourse connectives are words or phrases that signal the presence of a discourse relation. |
| Approach: | They propose a model that generates discourse connectives between arguments and predicts discourse relations based on the generated connectives. |
| Outcome: | The proposed model outperforms baselines on three datasets and is highly accurate. |
Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various Domains (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and alternative lexicalizations. |
| Approach: | They propose a model for extracting and classifying discourse relation signals from the Penn Discourse Treebank v3 corpus and introduce a new way of modeling rhetorical style by the linear order of coherence relations. |
| Outcome: | The proposed models are based on the Penn Discourse Treebank v3 corpus and employ n-gram patterns to predict genre/domain discrimination. |