Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.

Similar Papers

Intertextual Correspondence for Integrating Corpora (L18-1)

Copied to clipboard

Challenge: Using intertextual correspondence, we can combine annotated text corpora to create new annotation connections.
Approach: They propose to use intertextual correspondence as an integrative technique for combining annotated text corpora.
Outcome: The proposed technique can be used to build argumentative arguments in two annotated text corpora.
An Assessment of Explicit Inter- and Intra-sentential Discourse Connectives in Turkish Discourse Bank (L18-1)

Copied to clipboard

Challenge: Discourse parsing is a challenging task for NLP.
Approach: They propose to add a new set of explicit intra-sentential connectives to Turkish Discourse Bank 1.1 . they propose to evaluate the converb sense annotations and compare them to other Turkish corpus .
Outcome: The proposed annotations show that the subordinators tend to select certain senses not selected by explicit inter- and intra-sentential discourse connectives in the data.
Persian Discourse Treebank and coreference corpus (L18-1)

Copied to clipboard

Challenge: Currently, we are adding a new document-level discourse annotation to our new corpus.
Approach: They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution.
Outcome: The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens.
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task.
Approach: They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank.
Outcome: The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank.
Adapting BERT to Implicit Discourse Relation Classification with a Focus on Discourse Connectives (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on the performance of BERT for implicit discourse relation classification have not been conducted.
Approach: They propose to apply BERT to implicit discourse relation classification by performing additional pre-training on text tailored to discourse relations.
Outcome: The proposed methods outperform previous state-of-the-art models in many tasks.
Is Partial Linguistic Information Sufficient for Discourse Connective Disambiguation? A Case Study of Concession (2025.acl-srw)

Copied to clipboard

Challenge: Discourse relations are often not linguistically marked, but there are various connectives that explicitly signal discourse relations.
Approach: They analyze linguistic features that play an important role in disambiguation of polysemous connectives in Japanese by performing a neural language model.
Outcome: The proposed model performed well after removal of one of the two arguments that constitute the discourse relation, but significantly degraded disambiguation performance.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various Domains (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and alternative lexicalizations.
Approach: They propose a model for extracting and classifying discourse relation signals from the Penn Discourse Treebank v3 corpus and introduce a new way of modeling rhetorical style by the linear order of coherence relations.
Outcome: The proposed models are based on the Penn Discourse Treebank v3 corpus and employ n-gram patterns to predict genre/domain discrimination.
Clarifying Underspecified Discourse Relations in Instructional Texts (2025.findings-acl)

Copied to clipboard

Challenge: Discourse relations can be optionally realized through explicit connectives such as “but” and “while”.
Approach: They build a corpus of 4,274 text revisions in which a connective was explicitly inserted . they collect plausibility annotations on other connectives to check whether they represent suitable alternatives .
Outcome: The proposed model predicts plausibility of individual connectives with up to 66% accuracy, but is not reliable when multiple relations are plausible.
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)

Copied to clipboard

Challenge: Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data.
Approach: They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes.
Outcome: The proposed corpus includes 112 argumentative microtexts and a new annotation tool.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations