Challenge: Recent work at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance .
Approach: They propose to characterize what constitutes an explanation that is itself "fair" they use not just accuracy and label time, but psychological impact of explanations on different groups .
Outcome: The proposed method is based on content moderation of potential hate speech and its differential impact on Asian vs. non-Asian proxy moderators across explanation approaches.

Similar Papers

Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster (2024.acl-short)

Copied to clipboard

Challenge: Existing studies have shown that explanations can support content moderators to make faster decisions, but the benefits of such models have not been studied.
Approach: They propose to use structured explanations to support content moderators to make faster decisions by 7.4%.
Outcome: The proposed models lower the speed of real-world moderators by 7.4% compared to generic explanations and are often ignored . previous studies have shown that explanations can support moderator's decision making by detecting violations of policies but the benefits have not been studied .
On the Interaction of Belief Bias and Explanations (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts .
Approach: They propose to account for belief bias in explainability by using models of varying quality and adversarial examples.
Outcome: The proposed methods show that results change when using models of varying quality and adversarial examples.
HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content (2026.findings-eacl)

Copied to clipboard

Challenge: Existing reward models for explaining hate speech are optimized for broad notions of safety, but they assign lower scores to contextually rich explanations.
Approach: They propose a reward model that integrates interpretable signals to better align reward scores with the needs of hate speech explanation.
Outcome: The proposed model outperforms general-purpose baselines and improves pair-wise preference.
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains (2025.findings-naacl)

Copied to clipboard

Challenge: a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech .
Approach: They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection .
Outcome: The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture .
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair .
Approach: They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness.
Outcome: The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms .
COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable Recommendation (2023.emnlp-main)

Copied to clipboard

Challenge: Personalized text generation (PTG) is a key component of our digital lives but can inadvertently associate different levels of linguistic quality with users’ protected attributes.
Approach: They propose a framework to achieve measure-specific counterfactual fairness in explanation generation by focusing on one of the most studied settings: generating natural language explanations for recommendations.
Outcome: The proposed framework achieves measure-specific counterfactual fairness in explanation generation.
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation (2026.acl-long)

Copied to clipboard

Challenge: Existing studies in Bangla focus on hate classification while overlooking interpretability.
Approach: They propose to create a dataset with human-annotated labels for banla that contains 19,203 YouTube comments spanning April 2024–June 2025.
Outcome: The proposed dataset outperforms existing datasets on open and closed-source LLMs on interpretability and better understanding of hate speech in linguistically rich yet under-resourced languages.
InterFair: Debiasing with Natural Language Feedback for Fair Interpretable Predictions (2023.emnlp-main)

Copied to clipboard

Challenge: Debiasing methods in NLP models focus on isolating information related to a sensitive attribute (e.g., gender or race) but instead argue that a favorable debiaser should use sensitive information ‘fairly,’ with explanations, rather than blindly eliminating it.
Approach: They propose that a favorable debiasing method should use sensitive information ‘fairly,’ with explanations, rather than blindly eliminating it.
Outcome: The proposed approach reduces bias in explanations while maintaining the same prediction accuracy.
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation.
Approach: They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language.
Outcome: The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters.
BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social Biases (2023.emnlp-main)

Copied to clipboard

Challenge: Toxicity annotators and content moderators often default to mental shortcuts when making decisions, leading to subtle toxicity being missed and seemingly harmless content being over-detected.
Approach: They propose a framework that provides AI-generated explanations of statements’ implied social biases to enhance content moderation setups.
Outcome: The proposed framework significantly improves content moderation setups by enabling users to think more thoroughly about their decisions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations