MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification (2024.naacl-long)
Copied to clipboard
| Challenge: | Fallacy classification is a task of broad importance due to advances in deep learning and availability of more data. |
| Approach: | They propose a new annotation scheme tailored for subjective NLP tasks and a method designed to handle subjectivity. |
| Outcome: | The proposed approach integrates existing fallacy classification datasets with new ones. |
Similar Papers
Fine-grained Fallacy Detection with Human Label Variation (2025.naacl-long)
Copied to clipboard
| Challenge: | Fallacy detection is an open challenge in NLP and has shown to be intrinsically difficult for both humans and machines. |
| Approach: | They propose a framework that minimizes annotation errors whilst keeping signals of human label variation. |
| Outcome: | The proposed framework minimizes annotation errors while keeping signals of human label variation. |
Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling (2024.emnlp-main)
Copied to clipboard
Irfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue, Camélia Guerraoui, Wenzhi Wang, Shoichi Naito, Jungmin Choi, Kentaro Inui
| Challenge: | Prior work on quality assessment has focused on numerical scoring and fallacy type-labeling tasks, without aiming to analyze fallacy logic structures. |
| Approach: | They propose four sets of explainable templates for common informal logical fallacies designed to explicate a fallacy’s implicit logic. |
| Outcome: | The proposed models achieve a high agreement score and reasonable coverage 83% on 400 fallacious arguments and state-of-the-art language models struggle with detecting fallacy templates (0.47 accuracy). |
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing work on fallacy recognition is still in its early stages, with limited datasets available. |
| Approach: | They propose to use GPT3.5 to generate synthetic examples and explore prompt settings to improve the representation of the infrequent classes. |
| Outcome: | The proposed model improves on existing models and generates synthetic examples with GPT3.5. |
CRASS: A Novel Data Set and Benchmark to Test Counterfactual Reasoning of Large Language Models (2022.lrec-1)
Copied to clipboard
| Challenge: | CRASS data set and benchmark provide novel test scheme to evaluate large language models . authors present and explain the CRAS data set, a novel basis to test reasoning and natural language understanding of LLMs . |
| Approach: | They introduce a new test scheme utilizing questionized counterfactual conditionals to evaluate large language models. |
| Outcome: | The proposed model sets out to be the most powerful and valid tool to evaluate large language models. |
The Search for Agreement on Logical Fallacy Annotation of an Infodemic (2022.lrec-1)
Copied to clipboard
Claire Bonial, Austin Blodgett, Taylor Hudson, Stephanie M. Lukin, Jeffrey Micher, Douglas Summers-Stay, Peter Sutor, Clare Voss
| Challenge: | a parallel "infodemic" has emerged with the COVID-19 pandemic . logical fallacies can be subtly encoded in the structure of a document across multiple sentences . |
| Approach: | They evaluate an annotation schema for labeling logical fallacy types using linguist annotations . they propose to use a machine learning algorithm to train annotators for fallacy detection . |
| Outcome: | The proposed annotation schema is clear and non-overlapping for manual and system assignment. |
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)
Copied to clipboard
| Challenge: | The first workshop on crowdsourcing for NLP is open to all . |
| Approach: | The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks. |
| Outcome: | The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data . |
Multimodal Fallacy Classification in Political Debates (2024.eacl-short)
Copied to clipboard
| Challenge: | Recent advances in NLP suggest that some tasks, such as argument detection and relation classification, are better framed in a multimodal perspective. |
| Approach: | They propose to use multimodal argument mining to capture paralinguistic aspects of fallacious arguments. |
| Outcome: | The proposed multimodal argument mining improves argument detection and relation classification in political debates. |
Adversarial NLI: A New Benchmark for Natural Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | a new large-scale NLI benchmark dataset is presented to test models on a variety of popular NLIs. |
| Approach: | They propose a large-scale NLI benchmark dataset that is iteratively compared with a human-and-model-in-the-loop procedure. |
| Outcome: | The proposed method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate. |
Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work has focused on the ability of Large Language Models to **identify** or **classify** fallacies, but their robustness against these fallacias in persuasive contexts remains largely unexplored. |
| Approach: | They propose a new metric to assess LLM robustness against fallacies by pairing factual questions with fallacious arguments and developing a multi-round debate framework to assess model resilience. |
| Outcome: | The proposed metric disentangles robustness from a model’s knowledge limitations and demonstrates unique vulnerability profiles across models. |
Argument-based Detection and Classification of Fallacies in Political Debates (2023.emnlp-main)
Copied to clipboard
| Challenge: | Fallacies are arguments that employ faulty reasoning, causing inaccurate conclusions and invalid inferences . ad hominem fallacy is one of the most common fallacy labels used in political debates despite its use in many scenarios . |
| Approach: | They extend the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments by incorporating the most recent Trump-Biden debate. |
| Outcome: | The proposed method extends the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments . |