AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts (2025.emnlp-main)
Copied to clipboard
| Challenge: | Distinguishing LLM-generated text from human-written is a key challenge for safe and ethical NLP, especially in high-stake settings such as persuasive online discourse. |
| Approach: | They propose to use general-purpose linguistic features and domain-specific features related to argument quality to compare human- and LLM-authored arguments. |
| Outcome: | The proposed framework compares arguments by humans and three LLMs using two easily-interpretable feature sets. |
Similar Papers
Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can generate highly persuasive text, raising concerns about misuse for propaganda, manipulation, and other harmful purposes. |
| Approach: | They propose a multilingual benchmark to compare LLM-generated persuasive texts with human-written ones. |
| Outcome: | The proposed benchmark compares human-authored and LLM-generated persuasive texts . it finds that overtly persuasive LLMs are easier to detect than human-written ones . |
“I understand your perspective”: LLM Persuasion through the Lens of Communicative Action Theory (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actions remains largely unexplored. |
| Approach: | They examine whether Large Language Models express illocutionary intent in ways comparable to human communication by simulated online discussions . |
| Outcome: | The proposed models express illocutionary intents in ways comparable to human communication, and crowd-sourced workers prefer them over human-written ones. |
Can Language Models Recognize Convincing Arguments? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have found that large language models can generate persuasive content without engaging in human experimentation. |
| Approach: | They extend a dataset with debates, votes, and user traits to measure LLMs' ability to distinguish between strong and weak arguments, predict stances based on beliefs and demographic characteristics, and determine appeal of argument to individual based upon their traits. |
| Outcome: | The proposed tasks outperform human predictions in detecting convincing arguments in debates, votes, and user traits. |
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings. |
| Approach: | They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations. |
| Outcome: | The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts. |
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that large language models can successfully persuade humans and amplify persuasive language. |
| Approach: | They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language. |
| Outcome: | The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified. |
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have focused on specific domains or types of persuasion, but a general study has focused on how LLMs produce persuasive text. |
| Approach: | They construct a dataset to measure and benchmark the ability of Large Language Models (LLMs) to produce persuasive text. |
| Outcome: | The proposed model can be used to generate persuasive text across domains and domains. |
Argument-Based Consistency in Toxicity Explanations of LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods to evaluate free-form toxicity explanations are overly relying on input text perturbations. |
| Approach: | They propose a multi-dimensional criterion to evaluate LLMs' reasoning about toxicity . they conduct experiments on three Llama models and an 8B Ministral model . |
| Outcome: | The proposed criterion measures the extent to which LLMs’ free-form toxicity explanations reflect an ideal and logical argumentation process. |
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks (2026.findings-acl)
Copied to clipboard
| Challenge: | Argumentation skills are an essential toolkit for large language models (LLMs). |
| Approach: | They propose a benchmark to evaluate the generalizability of five LLM families across 46 computational argumentation tasks. |
| Outcome: | The proposed benchmark evaluates the generalizability of five LLM families across 46 computational argumentation tasks covering mining arguments, assessing perspectives, evaluating argument quality, reasoning about arguments, and generating arguments. |
Exploring the Potential of Large Language Models in Computational Argumentation (2024.acl-long)
Copied to clipboard
| Challenge: | Argumentation is an essential tool in various domains, including law, public policy, and artificial intelligence. |
| Approach: | They propose to evaluate LLMs on various computational argumentation tasks . they organize existing tasks into six main categories and standardize the format of 14 datasets . |
| Outcome: | The proposed model performs well on argument mining and argument generation tasks. |
Can Large Language Models Understand Argument Schemes? (2025.findings-acl)
Copied to clipboard
| Challenge: | Argument schemes are stereotypical forms of reasoning that occur in everyday arguments. |
| Approach: | They propose to use large language models (LLMs) to classify argument schemes based on Walton’s taxonomy to employ formal definitions and LLM-generated descriptions to enhance task instructions. |
| Outcome: | The proposed models perform well on annotated and automatically generated arguments, and provide insights for advancing reasoning capabilities in computational argumentation. |