Challenge: a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index .
Approach: They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries .
Outcome: The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english .

Similar Papers

Scaling up Discourse Quality Annotation for Political Science (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotations on deliberative quality are time-consuming and suffer from class imbalance . ephd thesis: deliberation is not only the output of the decision making, but also the discussion that leads up to it.
Approach: They propose to use data augmentation techniques to improve deliberative quality predictions in a standard dataset.
Outcome: The proposed methods outperform classifiers based on linguistic features and argument quality annotations with or without data augmentation.
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are ubiquitous in modern NLP, but ethical questions have been raised about their use as analysis tools.
Approach: They propose a framework that transforms noisy, multi-topic contributions into argumentative units ready for downstream analysis.
Outcome: The proposed framework can be run locally and transparently with limited resources.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
Computational Analysis of Political Texts: Bridging Research Efforts Across Communities (P19-4)

Copied to clipboard

Challenge: Political scientists have developed and adopted natural language processing (NLP) methods to exploit text as an additional source of data in their analyses.
Approach: This tutorial aims to provide a gentle introduction to methods and tasks related to computational analysis of political texts from both communities.
Outcome: The main goal of this tutorial is to bring the two research communities closer to each other and contribute to faster and more significant developments in this interdisciplinary area.
FREDSum: A Dialogue Summarization Corpus for French Political Debates (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in deep learning have improved the performance of abstractive summarization systems.
Approach: They present a dataset of french political debates to enhance resources for multi-lingual dialogue summarization.
Outcome: The proposed dataset will be made publicly available for use by the research community.
A Corpus for Modeling User and Language Effects in Argumentation on Online Debating (P19-1)

Copied to clipboard

Challenge: Existing argumentation datasets have allowed only limited assessment of "user" traits because information on background of users is generally unavailable.
Approach: They present a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles.
Outcome: The proposed dataset includes 78,376 debates generated over a 10-year period along with comprehensive participant profiles.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
Let’s discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment (2024.emnlp-main)

Copied to clipboard

Challenge: Argumentation is a key competence and an important cultural technique in democratic societies.
Approach: They propose to create domain-specific datasets and methods to assess argument quality.
Outcome: The proposed methods address gaps in the literature and aid future research in the domain.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
An Annotated Corpus for Sexism Detection in French Tweets (2020.lrec-1)

Copied to clipboard

Challenge: Social media networks allow users to share opinions and sentiments, which can cause a large spreading of hatred or abusive messages.
Approach: They propose to annotate 12,000 tweets with a sexism detection scheme in France . they propose to use deep learning to detect if a message with sexist content is really s.
Outcome: The proposed scheme detects sexist content and identifies if it is really sexism . the proposed scheme is the first of its kind in the u.s.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations