Developing A Multilabel Corpus for the Quality Assessment of Online Political Talk (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index . |
| Approach: | They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries . |
| Outcome: | The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english . |
Similar Papers
Scaling up Discourse Quality Annotation for Political Science (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing annotations on deliberative quality are time-consuming and suffer from class imbalance . ephd thesis: deliberation is not only the output of the decision making, but also the discussion that leads up to it. |
| Approach: | They propose to use data augmentation techniques to improve deliberative quality predictions in a standard dataset. |
| Outcome: | The proposed methods outperform classifiers based on linguistic features and argument quality annotations with or without data augmentation. |
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are ubiquitous in modern NLP, but ethical questions have been raised about their use as analysis tools. |
| Approach: | They propose a framework that transforms noisy, multi-topic contributions into argumentative units ready for downstream analysis. |
| Outcome: | The proposed framework can be run locally and transparently with limited resources. |
An Environment for Relational Annotation of Political Debates (P19-3)
Copied to clipboard
| Challenge: | Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities. |
| Approach: | They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors. |
| Outcome: | The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors. |
Computational Analysis of Political Texts: Bridging Research Efforts Across Communities (P19-4)
Copied to clipboard
| Challenge: | Political scientists have developed and adopted natural language processing (NLP) methods to exploit text as an additional source of data in their analyses. |
| Approach: | This tutorial aims to provide a gentle introduction to methods and tasks related to computational analysis of political texts from both communities. |
| Outcome: | The main goal of this tutorial is to bring the two research communities closer to each other and contribute to faster and more significant developments in this interdisciplinary area. |
FREDSum: A Dialogue Summarization Corpus for French Political Debates (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in deep learning have improved the performance of abstractive summarization systems. |
| Approach: | They present a dataset of french political debates to enhance resources for multi-lingual dialogue summarization. |
| Outcome: | The proposed dataset will be made publicly available for use by the research community. |
A Corpus for Modeling User and Language Effects in Argumentation on Online Debating (P19-1)
Copied to clipboard
| Challenge: | Existing argumentation datasets have allowed only limited assessment of "user" traits because information on background of users is generally unavailable. |
| Approach: | They present a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles. |
| Outcome: | The proposed dataset includes 78,376 debates generated over a 10-year period along with comprehensive participant profiles. |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |
Let’s discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Argumentation is a key competence and an important cultural technique in democratic societies. |
| Approach: | They propose to create domain-specific datasets and methods to assess argument quality. |
| Outcome: | The proposed methods address gaps in the literature and aid future research in the domain. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
An Annotated Corpus for Sexism Detection in French Tweets (2020.lrec-1)
Copied to clipboard
Patricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari, Gloria Origgi, Marlène Coulomb-Gully
| Challenge: | Social media networks allow users to share opinions and sentiments, which can cause a large spreading of hatred or abusive messages. |
| Approach: | They propose to annotate 12,000 tweets with a sexism detection scheme in France . they propose to use deep learning to detect if a message with sexist content is really s. |
| Outcome: | The proposed scheme detects sexist content and identifies if it is really sexism . the proposed scheme is the first of its kind in the u.s. |