Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing static benchmarks for harmful content detection face limitations in scalability and diversity. |
| Approach: | They propose a framework for synthesizing harmful content using persona-guided large language model agents. |
| Outcome: | The proposed framework achieves a high success rate in harmful generation tests across multiple detection systems. |
Similar Papers
Synthia: Scalable Grounded Persona Generation from Social Media Data (2026.acl-long)
Copied to clipboard
| Challenge: | Persona-driven large language models (LLMs) are increasingly used in computational social science, yet their validity critically depends on the fidelity of the underlying personas. |
| Approach: | They propose a persona-generation framework that grounds LLM-generated personas in real social-media posts while delegating narrative construction to language models. |
| Outcome: | The proposed framework outperforms state-of-the-art methods for most demographics across different dimensions while maintaining interaction graph structure among personas grounded in real social network users. |
Detoxifying Large Language Models via the Diversity of Toxic Samples (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for analyzing and utilizing toxic samples are limited . current methods fail to fully harness their potential . |
| Approach: | They propose a diverse detoxification framework that leverages toxic samples' diversity . they propose MPSG strategy and SC-DPO approach to elicit personalized toxic responses . |
| Outcome: | The proposed framework could be used to optimize large language models for user safety . it incorporates two components: MPSG strategy and SC-DPO approach . |
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on self-reported user data or interviews . a recent study evaluated how Replika responds to high-risk user groups . |
| Approach: | They propose a framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications. |
| Outcome: | The proposed framework evaluates how Replika responds to high-risk user groups . it incorporates emotion modeling and LLM-assisted utterance-and harm-level classification . |
Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey (2023.eacl-main)
Copied to clipboard
| Challenge: | Recent advances in the capacity of large language models to generate human-like text have prompted a heated discourse around the risks of societal harms they introduce. |
| Approach: | They propose a taxonomy of interventions organized around the different phases where they can be adopted to mitigate harms. |
| Outcome: | The proposed methods are based on several prior works’ taxonomies of language model risks and provide an overview of strategies for detecting and ameliorating different kinds of risks/harms. |
Mitigating Societal Harms in Large Language Models (2023.emnlp-tutorial)
Copied to clipboard
| Challenge: | Recent studies have highlighted societal harms that can be caused by language generation models deployed in the wild. |
| Approach: | They propose to use a typology of technical approaches to mitigating harms of language generation models to provide an overview of potential social issues in language generation including toxicity, social biases, misinformation, factual inconsistency, and privacy violations. |
| Outcome: | The proposed typology addresses toxicity, biases, misinformation, factual inconsistency, and privacy violations in language generation models. |
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for detecting harmful content lack diversity and quality of datasets. |
| Approach: | They propose a framework for synthesizing toxic information from social media datasets . their framework generates a wide variety of synthetic, yet remarkably realistic, examples of toxic information . |
| Outcome: | The proposed framework can generate a wide variety of synthetic, yet remarkably realistic, examples of toxic information. |
Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models enable them to follow freeform instructions, including imitating generic or specific demographic personas in conversations. |
| Approach: | They propose to investigate persona biases by experimenting with UNIVERSALPERSONA, a model that incorporates both generic and specific personas. |
| Outcome: | The proposed model systematically measures persona biases in harmful expression and harmful agreement. |
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)
Copied to clipboard
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Smith
| Challenge: | generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say . |
| Approach: | They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |
| Outcome: | The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |
USB: A COMPREHENSIVE AND UNIFIED SAFETY EVALUATION BENCHMARK FOR MULTIMODAL LARGE LANGUAGE MODELS (2026.acl-long)
Copied to clipboard
Baolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang
| Challenge: | Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations. |
| Approach: | They propose a framework that covers 61 risk categories across four modality interactions to address this gap. |
| Outcome: | The proposed framework covers 61 risk categories across four distinct modality interactions. |
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to generate toxic content by large language models are based on pipelines . current approaches focus on preserving performance while effectively mitigating toxicity . |
| Approach: | They propose a framework for implicit knowledge editing and controlled text generation by using hard negatives. |
| Outcome: | The proposed framework significantly reduces toxic generation while maintaining strong performance on downstream tasks. |