Towards Conceptualization of “Fair Explanation”: Disparate Impacts of anti-Asian Hate Speech Explanations on Content Moderators (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent work at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance . |
| Approach: | They propose to characterize what constitutes an explanation that is itself "fair" they use not just accuracy and label time, but psychological impact of explanations on different groups . |
| Outcome: | The proposed method is based on content moderation of potential hate speech and its differential impact on Asian vs. non-Asian proxy moderators across explanation approaches. |
Similar Papers
Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster (2024.acl-short)
Copied to clipboard
Agostina Calabrese, Leonardo Neves, Neil Shah, Maarten Bos, Björn Ross, Mirella Lapata, Francesco Barbieri
| Challenge: | Existing studies have shown that explanations can support content moderators to make faster decisions, but the benefits of such models have not been studied. |
| Approach: | They propose to use structured explanations to support content moderators to make faster decisions by 7.4%. |
| Outcome: | The proposed models lower the speed of real-world moderators by 7.4% compared to generic explanations and are often ignored . previous studies have shown that explanations can support moderator's decision making by detecting violations of policies but the benefits have not been studied . |
On the Interaction of Belief Bias and Explanations (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts . |
| Approach: | They propose to account for belief bias in explainability by using models of varying quality and adversarial examples. |
| Outcome: | The proposed methods show that results change when using models of varying quality and adversarial examples. |
HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing reward models for explaining hate speech are optimized for broad notions of safety, but they assign lower scores to contextually rich explanations. |
| Approach: | They propose a reward model that integrates interpretable signals to better align reward scores with the needs of hate speech explanation. |
| Outcome: | The proposed model outperforms general-purpose baselines and improves pair-wise preference. |
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains (2025.findings-naacl)
Copied to clipboard
| Challenge: | a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech . |
| Approach: | They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection . |
| Outcome: | The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture . |
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair . |
| Approach: | They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness. |
| Outcome: | The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms . |
COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable Recommendation (2023.emnlp-main)
Copied to clipboard
Nan Wang, Qifan Wang, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, Shaoliang Nie
| Challenge: | Personalized text generation (PTG) is a key component of our digital lives but can inadvertently associate different levels of linguistic quality with users’ protected attributes. |
| Approach: | They propose a framework to achieve measure-specific counterfactual fairness in explanation generation by focusing on one of the most studied settings: generating natural language explanations for recommendations. |
| Outcome: | The proposed framework achieves measure-specific counterfactual fairness in explanation generation. |
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation (2026.acl-long)
Copied to clipboard
Faisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md Tahmid Hasan Fuad, Md Farhan Ishmam, Faria Sultana, M Ashraful Amin, Amin Ahsan Ali, Akmmahbubur Rahman
| Challenge: | Existing studies in Bangla focus on hate classification while overlooking interpretability. |
| Approach: | They propose to create a dataset with human-annotated labels for banla that contains 19,203 YouTube comments spanning April 2024–June 2025. |
| Outcome: | The proposed dataset outperforms existing datasets on open and closed-source LLMs on interpretability and better understanding of hate speech in linguistically rich yet under-resourced languages. |
InterFair: Debiasing with Natural Language Feedback for Fair Interpretable Predictions (2023.emnlp-main)
Copied to clipboard
| Challenge: | Debiasing methods in NLP models focus on isolating information related to a sensitive attribute (e.g., gender or race) but instead argue that a favorable debiaser should use sensitive information ‘fairly,’ with explanations, rather than blindly eliminating it. |
| Approach: | They propose that a favorable debiasing method should use sensitive information ‘fairly,’ with explanations, rather than blindly eliminating it. |
| Outcome: | The proposed approach reduces bias in explanations while maintaining the same prediction accuracy. |
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation. |
| Approach: | They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language. |
| Outcome: | The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters. |
BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social Biases (2023.emnlp-main)
Copied to clipboard
| Challenge: | Toxicity annotators and content moderators often default to mental shortcuts when making decisions, leading to subtle toxicity being missed and seemingly harmless content being over-detected. |
| Approach: | They propose a framework that provides AI-generated explanations of statements’ implied social biases to enhance content moderation setups. |
| Outcome: | The proposed framework significantly improves content moderation setups by enabling users to think more thoroughly about their decisions. |