Challenge: DiscoGeM is a crowdsourced corpus of 6,505 implicit discourse relations . the results show that a significant proportion of discourse relations are ambiguous . text genre is crucially affected by the distribution of discourse relation labels .
Approach: They propose to use crowdsourced corpus of 6,505 implicit discourse relations to classify relations . they propose to include genre as a factor in automatic relation classification .
Outcome: The proposed dataset shows that a significant proportion of discourse relations are ambiguous and can express multiple relation senses.

Similar Papers

DiscoGeM 2.0: A Parallel Corpus of English, German, French and Czech Implicit Discourse Relations (2024.lrec-main)

Copied to clipboard

Challenge: DiscoGeM 2.0 is a crowdsourced, parallel corpus of 12,834 implicit discourse relations . implicit discourse relationships are highly ambiguous and can have various interpretations .
Approach: They propose a crowdsourced annotation method that can be extended to other languages . they propose to annotate 12,834 implicit discourse relations in German, German, French and Czech data .
Outcome: The proposed method can be extended to other languages and reveals that implicit relations inferred in one language may differ from those inferted in the translation.
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training (2022.lrec-1)

Copied to clipboard

Challenge: Recent methods have obtained promising results by extracting relation labels from participants . obtaining linguistic annotations from novice crowdworkers is difficult . crowdsourcing allows for fast and cost-effective collection of labelled data, but because tasks need to be intuitive, crowdworker cannot be asked to perform them.
Approach: They propose to use a selection-only approach to obtain linguistic annotations from novices . current study shows that the method is cost- and time-intensive .
Outcome: The current study shows that selection and training improves the agreement between workers and gold labels, but the method is cost- and time-intensive.
Multi-Label Classification for Implicit Discourse Relation Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Prior research in discourse relation recognition has treated these instances as separate examples during training, with a gold-standard prediction matching one of the labels considered correct at test time.
Approach: They propose to use multiple labels to annotate an example when multiple relations are believed to hold simultaneously.
Outcome: The proposed frameworks don't depress performance for single-label prediction.
Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design (2023.tacl-1)

Copied to clipboard

Challenge: Disagreement in natural language annotation has been studied from a perspective of biases introduced by the annotators and the annotation frameworks.
Approach: They propose to analyze task design bias in crowdsourced annotations where lay annotators are used to elicit interpretations.
Outcome: The proposed methods can push annotators towards certain relations and some discourse relation senses can be better elicited with one or the other approach.
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature.
Approach: They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders .
Outcome: The proposed evaluation protocol improves the existing framework and provides strong baseline results.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)

Copied to clipboard

Challenge: Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old.
Approach: They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing.
Outcome: The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain.
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing relation extraction methods rely on exact matching with human-annotated reference relations, while GRE methods produce diverse and semantically accurate relations.
Approach: They propose a multi-dimensional assessment of relation extraction methods using human-annotated reference relations.
Outcome: The proposed method is consistent with human preferences for RE quality.
QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines (2020.emnlp-main)

Copied to clipboard

Challenge: Discourse relations describe how two propositions relate to one another . annotating discourse relations requires expert annotators .
Approach: They propose a new representation of discourse relations as question-and-answer pairs that crowd-sources wide-coverage data annotated with discourse relations.
Outcome: The proposed representation of discourse relations as QA pairs allows crowd-sourcing wide-coverage datasets annotated with discourse relations.
A Streamlined Method for Sourcing Discourse-level Argumentation Annotations from the Crowd (N19-1)

Copied to clipboard

Challenge: Existing methods for analyzing discourse-level argument annotations require expensive labor and data.
Approach: They propose a method that breaks down a popular but complex discourse-level argument annotation scheme into a simple iterative procedure that can be applied even by untrained annotators.
Outcome: The proposed method can be applied even by untrained annotators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations