Written Justifications are Key to Aggregate Crowdsourced Forecasts (2021.findings-emnlp)

Copied to clipboard

Challenge: aggregating crowdsourced forecasts benefits from modeling written justifications . a majority of respondents support the idea that crowds are more reliable than experts .
Approach: They propose to model written justifications for crowdsourced questions by analyzing their results in a literature review.
Outcome: The results show that the written justifications are beneficial to call a question throughout its life except in the last quarter.

Similar Papers

FORECAST2023: A Forecast and Reasoning Corpus of Argumentation Structures (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on the role of reasoning in forecasting has focused on surface-level features such as linguistic markers, the use of comparison classes, and overall dialectical complexity.
Approach: They propose to use a dataset of such prediction rationales to create a fully automated annotation system that can be used to enhance the argumentation.
Outcome: The proposed dataset provides a uniquely fine-grained and close characterisation of the structure of argumentation with potential impact on forecasting domains from intelligence analysis to investment decision-making.
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)

Copied to clipboard

Challenge: The first workshop on crowdsourcing for NLP is open to all .
Approach: The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks.
Outcome: The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data .
Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries (2022.naacl-main)

Copied to clipboard

Challenge: Existing pre-trained summarization models produce text that is factually inconsistent with the input.
Approach: They present a scale-based scale for Likert rating and a scoring algorithm for Best-Worst Scaling to improve crowdsourcing reliability.
Outcome: The proposed model is more reliable than existing models on two news summarization datasets.
Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language Inference (P19-1)

Copied to clipboard

Challenge: Recent advances on abstractive summarization have led to fluent summaries, but factual errors in generated summary still severely limit their use in practice.
Approach: They evaluate summaries produced by state-of-the-art models via crowdsourcing and show that factual errors occur frequently.
Outcome: The proposed models can detect errors and reduce them by reranking alternative summaries.
Event2Mind: Commonsense Inference on Events, Intents, and Reactions (P18-1)

Copied to clipboard

Challenge: Using a crowdsourced corpus of 25,000 event phrases, we construct a new task that uses commonsense reasoning to reason about the likely intents and reactions of the event participants.
Approach: They construct a crowdsourced corpus of 25,000 event phrases and use them to construct 'commonsense inference' they demonstrate that neural encoder-decoder models can compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants.
Outcome: The proposed task can be used to uncover implicit gender inequality in movie scripts.
A Dataset of Crowdsourced Word Sequences: Collections and Answer Aggregation for Ground Truth Creation (D19-59)

Copied to clipboard

Challenge: Existing work on answer aggregation for labels is limited . existing work on label aggregations is limited to label .
Approach: They propose three approaches to extractive word sequence aggregation from translated sentences generated by multiple workers.
Outcome: The proposed dataset contains translated sentences generated from multiple workers.
Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning (2023.acl-long)

Copied to clipboard

Challenge: Annotator disagreements are resolved before learning takes place, but researchers question the performance of a system when annotators disagree.
Approach: They propose a method that uses language features and label distributions to pool similar items into larger labels.
Outcome: The proposed method is based on five publicly available datasets with varying levels of disagreements on social media and in the wild using a dataset from Facebook.
An Empirical Study of Building a Strong Baseline for Constituency Parsing (P18-2)

Copied to clipboard

Challenge: Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization.
Approach: They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing.
Outcome: The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing.
Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects (D19-1)

Copied to clipboard

Challenge: Existing approaches to generating reviews struggle to generate justifications that are relevant to users’ decision-making process.
Approach: They propose an ‘extractive’ approach to identify review segments which justify users’ intentions and use it to distantly label massive review corpora and construct large-scale personalized recommendation justification datasets.
Outcome: The proposed model can generate convincing and diverse justifications from massive review corpora and distantly label massive review data.
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations