Papers by Benno Stein
Employing Argumentation Knowledge Graphs for Neural Argument Generation (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for generating arguments use end-to-end knowledge graphs or are controlled with respect to the argument's topic, aspects, or stance. |
| Approach: | They construct and populate three knowledge graphs and encode them into debate portals and relevant paragraphs from Wikipedia. |
| Outcome: | The proposed model produces arguments with superior quality than those generated without knowledge. |
Modeling Deliberative Argumentation Strategies on Wikipedia (P18-1)
Copied to clipboard
| Challenge: | Existing models for deliberative discussions have been built manually based on a small set of discussions, resulting in a level of abstraction that is not suitable for move recommendation. |
| Approach: | They propose to model argumentation strategies of deliberative discussions by annotating ongoing discussions with a label that can be used for move description. |
| Outcome: | The proposed model can predict arguments of participants in deliberative discussions using metadata from Wikipedia talk pages. |
Identifying the Human Values behind Arguments (2022.acl-long)
Copied to clipboard
| Challenge: | et al., 2003) examines human values in natural language arguments . authors provide a dataset of 5270 arguments from four geographical cultures . |
| Approach: | They propose a multi-level taxonomy of human values with 54 values and a dataset of 5270 arguments from four geographical cultures, manually annotated for human values. |
| Outcome: | The proposed model shows that human values are more diverse than previously thought . it shows that people disagree on the best course forward on controversial issues . |
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology . |
| Approach: | They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology . |
| Outcome: | The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors. |
Improving Argument Effectiveness Across Ideologies using Instruction-tuned Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a study finds that different political ideologies hold different worldviews, which leads to contentious debates . argument effectiveness is improved by using instruction-tuned large language models . |
| Approach: | They propose to use instruction-tuned large language models to turn ineffective arguments into effective arguments for people with certain ideologies. |
| Outcome: | The proposed methods improve argument effectiveness for liberals by rewriting arguments using three LLM methods. |
Heuristic Authorship Obfuscation (P19-1)
Copied to clipboard
| Challenge: | Existing methods for authorship verification are insufficient to control the authorial style of a text. |
| Approach: | They propose a novel method that models writing style difference as the Jensen-Shannon distance between character n-gram distributions of texts and manipulates an author’s subconsciously encoded writing style using heuristic search. |
| Outcome: | The proposed approach defeats state-of-the-art verification approaches while keeping text changes at a minimum. |
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants (N18-1)
Copied to clipboard
| Challenge: | Existing methods for analyzing warrants in natural language arguments are insufficient. |
| Approach: | They propose a method for reconstructing warrants from news comments . they use a crowdsourcing process to obtain warrants for 2k authentic arguments . |
| Outcome: | The proposed method will define a substantial step towards automatic warrant reconstruction. |
Bias Analysis and Mitigation in the Evaluation of Authorship Verification (P19-1)
Copied to clipboard
| Challenge: | a paper on authorship verification shows that the underlying experiment design cannot guarantee pushing forward the state of the art. |
| Approach: | They propose a "Basic and Fairly Flawed" authorship verifier that is on a par with the best approaches submitted so far . they pinpoint sources of bias that should be eliminated and propose 'refined' authorship corpus as effective countermeasure. |
| Outcome: | The proposed approach is on par with the best approaches submitted so far . the proposed approach shows that sources of bias should be eliminated . |
Paraphrase Acquisition from Image Captions (2023.eacl-main)
Copied to clipboard
| Challenge: | Using image captions, we hypothesize that different captions for the same image naturally form a set of mutual paraphrases. |
| Approach: | They propose to use image captions as a previously underutilized resource for paraphrases . they analyze captions in the English Wikipedia to find common paraphrase similarities . |
| Outcome: | The proposed dataset compares known paraphrase corpora with their syntactic and semantic similarity to the existing dataset. |
Detecting Media Bias in News Articles using Gaussian Bias Distributions (2020.findings-emnlp)
Copied to clipboard
| Challenge: | a new study shows that media bias is not only about honesty or accuracy, but also about taste or preference. |
| Approach: | They propose to use second-order information to detect media bias in articles . they propose to analyze the frequency, positions, and sequential order of biased statements . |
| Outcome: | The proposed model outperforms other models that use second-order information on biased statements on an existing media bias dataset. |
Controlled Neural Sentence-Level Reframing of News Articles (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a news article is framed from a specific perspective, but reframing can be difficult . a framed article can be used to communicate with opposing camps of audiences . |
| Approach: | They propose to reframe news articles using a media frame corpus to achieve this . they propose three strategies to train neural models for reframing . |
| Outcome: | The proposed techniques maintain coherence of sentences and reframe them correctly . the proposed techniques are effective but have tradeoffs . |
Argumentation and Domain Discourse in Scholarly Articles on the Theory of International Relations (2025.coling-main)
Copied to clipboard
Magdalena Wolska, Sassan Gholiagha, Mitja Sienknecht, Dora Kiesel, Irene Lopez Garcia, Patrick Riehmann, Matti Wiegmann, Bernd Froehlich, Katrin Girgensohn, Jürgen Neyer, Benno Stein
| Challenge: | SKILL project aims to provide students with AI tools to facilitate analysis of argumentation in scholarly articles on international relations. |
| Approach: | They propose to use AI to analyze argumentation in scholarly articles on international relations . they use a dataset, discourse analysis, and baseline experiments to examine argumentation and domain content types . |
| Outcome: | The proposed method enables educationally-relevant insight into scholarly IR discourse . it requires domain-specific training and fine-tuning on relation and content type prediction tasks. |
Before Name-Calling: Dynamics and Triggers of Ad Hominem Fallacies in Web Argumentation (N18-1)
Copied to clipboard
| Challenge: | Existing research lacks solid empirical investigation of typology of ad hominem arguments and their potential causes. |
| Approach: | They propose to perform several large-scale annotation studies and experiment with various neural architectures to validate hypotheses such as controversy or reasonableness. |
| Outcome: | The proposed model identifies the ad hominem fallacy and its possible causes using explainable neural network architectures. |
Analyzing the Persuasive Effect of Style in News Editorial Argumentation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing research has investigated the persuasive effect of content and style on argumentative content. |
| Approach: | They compare the style of news editorials with ideology-specific effect annotations to find out how important it is to achieve persuasiveness. |
| Outcome: | The proposed method shows that conservative readers are resistant to style on liberal editorials, whereas conservative readers resist style on conservatives. |
Mining Health-related Cause-Effect Statements with High Precision at Large Scale (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for assessing the health relatedness of phrases and sentences are slower and less effective than state-of-the-art medical entity linkers. |
| Approach: | They propose a termhood score that achieves 69% recall at over 90% precision on a web dataset with cause-effect statements. |
| Outcome: | The proposed method achieves 69% recall at over 90% precision on a web dataset with cause-effect statements. |
Unveiling the Power of Argument Arrangement in Online Persuasive Discussions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study shows that the CMV is the best time period in human history for the vast majority of people. |
| Approach: | They extend a semantic argumentation unit type model by clustering type sequences into different argument arrangement patterns and representing discussions as sequences of these patterns. |
| Outcome: | The proposed model outperforms existing classifiers on the change my view forum discussion data. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Trigger Warning Assignment as a Multi-Label Document Classification Problem (2023.acl-long)
Copied to clipboard
| Challenge: | a trigger warning is used to warn people about potentially disturbing content . a webis dataset of 1 million fanfiction works contains up to 36 different warnings per document . |
| Approach: | They introduce a multi-label task to assign a trigger warning to fanfiction . they map 41 million free-form tags assigned by authors into a single taxonomy of trigger warnings . |
| Outcome: | The proposed model achieves micro-F1 scores of about 0.5, which reveals the difficulty of the task. |
Differential Bias: On the Perceptibility of Stance Imbalance in Argumentation (2022.findings-aacl)
Copied to clipboard
| Challenge: | a theoretical model of bias classification is not feasible because of complexity of interpreting language phenomena. |
| Approach: | They propose to analyze whether a text is biased based on an algorithmic analysis . they propose to use a model to determine whether x is more biased than y . |
| Outcome: | a crowdsourcing study shows that differences in stance bias are perceptible when (light) support is provided through training or visual aids. |
Visualization of the Topic Space of Argument Search Results in args.me (D18-2)
Copied to clipboard
Yamen Ajjour, Henning Wachsmuth, Dora Kiesel, Patrick Riehmann, Fan Fan, Giuliano Castiglia, Rosemary Adejoh, Bernd Fröhlich, Benno Stein
| Challenge: | args.me is the first search engine for controversial topics . it ranks pro and con arguments by their relevance to a topic . |
| Approach: | They propose a visualization interface for result exploration that provides an overview of main aspects in a barycentric coordinate system. |
| Outcome: | The proposed search engine is the first dedicated argument search engine on the web. |
Unraveling the Search Space of Abusive Language in Wikipedia with Dynamic Lexicon Acquisition (D19-50)
Copied to clipboard
| Challenge: | Existing methods to detect abusive language only train one classifier for the whole variety of offending . a new method can support a moderator with explicit unraveled explanations for why something was flagged as abusive . |
| Approach: | a new method is proposed to distinguish explicitly abusive cases from the more "shadowed" ones . the researchers extend a lexicon of abusive terms to include new obfuscations of abusive words . |
| Outcome: | a new method can distinguish explicitly abusive cases from the more "shadowed" ones . the method can support a moderator with explicit unraveled explanations for why something was flagged as abusive . |
Reference-guided Style-Consistent Content Transfer (2024.lrec-main)
Copied to clipboard
| Challenge: | Text style transfer involves changing the style of a text while preserving its original style. |
| Approach: | They propose a task of style-consistent content transfer which involves modifying a text’s content based on a provided reference statement while preserving its original style. |
| Outcome: | The proposed approach meets three important conditions: reference faithfulness, style adherence, and coherence. |
Exploiting Personal Characteristics of Debaters for Predicting Persuasiveness (2020.acl-main)
Copied to clipboard
| Challenge: | Several studies have examined persuasiveness in debates by probing the main factors for establishing persuasion, particularly regarding the role of linguistic features of debaters' arguments. |
| Approach: | They propose to model debaters’ prior beliefs, interests, and personality traits based on their previous activity without dependence on explicit user profiles or questionnaires. |
| Outcome: | The proposed model improves persuasiveness prediction and debater resistance to persuasion. |
A Stylometric Inquiry into Hyperpartisan and Fake News (P18-1)
Copied to clipboard
| Challenge: | a style analysis of hyperpartisan news and fake news can distinguish them from mainstream news . left-wing and right-wing news share significantly more stylistic similarities than mainstream news does . |
| Approach: | a comparative style analysis of hyperpartisan news and fake news is carried out . authors show that left-wing and right-wing news share significantly more stylistic similarities . |
| Outcome: | a style analysis can distinguish hyperpartisan news from mainstream, satire from both . left-wing and right-wing news share significantly more stylistic similarities than mainstream . |
Celebrity Profiling (P19-1)
Copied to clipboard
| Challenge: | Using a corpus of 71,706 verified accounts, we construct a profile of a wide cross-section of local and global celebrities. |
| Approach: | They propose to use Twitter feeds of 71,706 verified accounts to build a corpus of celebrity profiles using Wikidata crawling. |
| Outcome: | The proposed corpus contains an average of 29,968 words per profile and up to 239 pieces of personal information. |
Topic Ontologies for Arguments (2023.findings-eacl)
Copied to clipboard
| Challenge: | Many computational argumentation tasks, such as stance classification, are topic-dependent. |
| Approach: | They map the argumentation landscape using the World Economic Forum, Wikipedia and Debatepedia as sources for argument topics. |
| Outcome: | The argument ontology is the first comprehensive assessment of argument topics in argument corpora. |
Crawling and Preprocessing Mailing Lists At Scale for Dialog Analysis (2020.acl-main)
Copied to clipboard
| Challenge: | a new neural segmentation model is used to segment 153 million emails . email is perhaps the most reliable and ubiquitous means of digital communication . |
| Approach: | They present a new neural segmentation model that crawls 153 million emails . it achieves 96% accuracy on 15 classes of email segments . |
| Outcome: | The proposed model achieves state-of-the-art performance while being more efficient to train than previous ones. |
Trigger Warnings: Bootstrapping a Violence Detector for Fan Fiction (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing guidelines for proactively alerting readers of potentially disturbing content have been proposed. |
| Approach: | They propose to use a labeled corpus of narrative fiction from a popular fan fiction site to determine whether to assign a trigger warning to an English story. |
| Outcome: | The proposed task achieves F1 scores between 0.8 and 0.9 on three datasets . the authors show that assigning trigger warnings for violence is feasible . |
Generalizing Unmasking for Short Texts (N19-1)
Copied to clipboard
| Challenge: | Authorship verification is the problem of inferring whether two texts were written by the same author. |
| Approach: | They propose a generalized unmasking approach which allows for authorship verification of short texts with high precision at an adjustable recall tradeoff. |
| Outcome: | The proposed approach achieves accuracies of 75–80% while allowing for easy adjustment to forensic scenarios that require higher levels of confidence. |
News Editorials: Towards Summarizing Long Argumentative Texts (2020.coling-main)
Copied to clipboard
| Challenge: | Using news summarization, we aim to target opinionated articles with a well-defined argumentation structure. |
| Approach: | They present a corpus of carefully curated summaries for 266 news editorials. |
| Outcome: | The summarization of opinionated articles with a well-defined argumentation structure is evaluated using a tailored annotation scheme. |
Crowdsourcing a Large Corpus of Clickbait on Twitter (C18-1)
Copied to clipboard
Martin Potthast, Tim Gollub, Kristof Komlossy, Sebastian Schuster, Matti Wiegmann, Erika Patricia Garces Fernandez, Matthias Hagen, Benno Stein
| Challenge: | Clickbait is a nuisance on social media. |
| Approach: | a corpus of 38,517 annotated Twitter tweets was constructed to detect clickbait . the corpus was annotating tweets on 4-point scale by five annotators at Amazon's Mechanical Turk . |
| Outcome: | The corpus of 38,517 annotated Twitter tweets was used to evaluate 12 clickbait detectors submitted to the Clickbait Challenge 2017 . |
Efficient Pairwise Annotation of Argument Quality (2020.acl-main)
Copied to clipboard
| Challenge: | Especially crowdsourcing suffers from assessors having different reference frames to base their judgments on and task instructions being nondescript and therefore unhelpful in ensuring consistency. |
| Approach: | They propose an efficient annotation framework for argument quality that uses a stochastic transitivity model and an effective sampling strategy to infer high-quality labels. |
| Outcome: | The proposed model significantly outperforms existing annotation procedures and offers statistical insights into argument quality. |
The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments (2024.lrec-main)
Copied to clipboard
Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Valentin Barriere, Doratossadat Dastgheib, Omid Ghahroodi, MohammadAli SadraeiJavaheri, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, Benno Stein
| Challenge: | Cultural norms can influence the prioritization of values, leading to distinct perspectives on debatable topics. |
| Approach: | They present a Touché23-ValueEval dataset that annotates 4780 new arguments and annotated 54 human values. |
| Outcome: | The Touché23-ValueEval dataset doubles the original Webis-ArgValués-22 dataset to 9324 arguments. |
Argumentation Synthesis following Rhetorical Strategies (C18-1)
Copied to clipboard
| Challenge: | Existing argument mining studies focus on logical structure of arguments, identifying their units and relations, and the effects of logical and emotional arguments across audiences. |
| Approach: | They propose to use rhetorical strategies to synthesize argumentative texts with different strategies. |
| Outcome: | The proposed model shows that the experts agree significantly more on selection when following the same strategy. |
CausalQA: A Benchmark for Causal Question Answering (2022.coling-1)
Copied to clipboard
Alexander Bondarenko, Magdalena Wolska, Stefan Heindorf, Lukas Blübaum, Axel-Cyrille Ngonga Ngomo, Benno Stein, Pavel Braslavski, Matthias Hagen, Martin Potthast
| Challenge: | Existing causal question answering datasets are relatively small and only include one type of causal question. |
| Approach: | They construct a benchmark corpus of 1.1 million causal questions with answers . they use a typology derived from a data-driven, manual analysis of QA datasets . |
| Outcome: | The proposed model achieves a ROUGE-L F1 score of 0.48 on the new QA benchmark. |
Analyzing Persuasion Strategies of Debaters on Social Media (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies on the analysis of persuasion in online discussions focus on the effectiveness of comments in individual discussions and ignore the effectiveness analysis of debaters over multiple discussions. |
| Approach: | They propose to quantify debaters effectiveness in the online discussion platform "ChangeMyView" they aim to explore diverse insights into their persuasion strategies . |
| Outcome: | The proposed analysis of debater effectiveness in the ChangeMyView subreddit reveals that debaters have different levels of effectiveness, behavioral characteristics and text stylistic features . |
Retrieval of the Best Counterargument without Prior Topic Knowledge (P18-1)
Copied to clipboard
| Challenge: | ad-hominem attacks are the most common form of argumentation in real life . |
| Approach: | They hypothesize that the best counterargument invokes the same aspects as the argument while having the opposite stance. |
| Outcome: | The proposed model is independent from the topic at hand, i.e., it applies to arbitrary arguments. |
Modeling Frames in Argumentation (D19-1)
Copied to clipboard
| Challenge: | In argumentation, framing is used to emphasize a specific aspect of a topic while concealing others. |
| Approach: | They propose an unsupervised method for framing arguments into non-overlapping frames . authors propose a corpus of 12, 326 debate-portal arguments organized along the frames of debates' topics . |
| Outcome: | The proposed method outperforms baselines on the argumentation task by 0.28 points. |
Task-Oriented Paraphrase Analytics (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on paraphrasing have applied different criteria to the task . authors have previously unmasked related tasks as paraphrases . |
| Approach: | They propose a taxonomy to organize 25 identified paraphrasing tasks . authors propose to use classifiers to identify tasks that a given paraphrased instance fits . |
| Outcome: | The proposed taxonomy identifies 25 paraphrasing tasks that fit the proposed task. |