Papers by Mona T. Diab
Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders (2026.acl-long)
Copied to clipboard
Xiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong, Zeyu Tang, Mona T. Diab, Virginia Smith, Kun Zhang
| Challenge: | Sparse Autoencoders (SAEs) are a tool in mechanistic interpretability (MI) but the aspiration to identify a canonical set of features is challenged by the observed inconsistency of learned SAE features across different training runs. |
| Approach: | They propose to use the Pairwise Dictionary Mean Correlation Coefficient to quantify SAE feature consistency as an evaluation axis alongside reconstruction and sparsity. |
| Outcome: | The proposed measure is based on the pairwise dictionary mean correlation coefficient (PW-MCC) on LLM activations. |
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models’ Detection of Human risky health behavior Content in Jirai Community (2026.eacl-long)
Copied to clipboard
Yunze Xiao, Tingyu He, Lionel Z. Wang, Yiming Ma, Xingyu Song, Xiaohang Xu, Mona T. Diab, Irene Li, Ka Chung Ng
| Challenge: | a cross-lingual dataset captures a transnational cultural phenomenon . risky health behaviors (RHB) are often linked to complex mental health conditions . |
| Approach: | They present the first cross-lingual dataset that captures a transnational cultural phenomenon . their dataset of more than 15,000 annotated social media posts forms the core of JiraiBench . |
| Outcome: | The study shows that cultural context can be more influential than linguistic similarity . the study also shows that the Japanese prompts better handle Chinese content . |
Personal Information Parroting in Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Modern language models memorize millions of PI instances, increasing privacy risks. |
| Approach: | They develop a model that parrots 13.6% of PI verbatim on a manually curated set of 483 instances . they recommend that pretraining datasets be aggressively filtered and anonymized to minimize PI parroting. |
| Outcome: | The proposed model outperforms the best regex-based PI detectors on a manually curated set of 483 instances of PI. |
Common Sense or Ableism? Rethinking Commonsense Reasoning Through the Lens of Disability (2026.eacl-short)
Copied to clipboard
| Challenge: | a recent study finds that commonsense reasoning is not always universal and can leave disabled people behind . a case study of disabled people with long COVID shows that common sense is not universal . |
| Approach: | They investigate how datasets and models deal with disability in commonsense reasoning . they use annotations from disabled and non-disabled persons for ableism . |
| Outcome: | The proposed datasets have low sensitivity to human-detected ableism but still detect 5 to 25% of entries as ableist. |
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs’ Cultural Processing of Figurative Language (2026.eacl-long)
Copied to clipboard
| Challenge: | Using figurative language as a proxy for cultural nuance and local knowledge, large language models struggle with connotative meaning. |
| Approach: | They evaluate large language models' ability to process culturally grounded language . they use figurative language as a proxy for cultural nuance and local knowledge . |
| Outcome: | The proposed models can understand and use figurative expressions that encode local knowledge and social nuance. |
PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers in Overleaf (2026.acl-demo)
Copied to clipboard
Jiarui Liu, Terry Jingchen Zhang, Ryan Faulkner, Xuanqiang Angelo Huang, Vilém Zouhar, Dominik Glandorf, Isabel Dahlgren, Rishit Dagli, Yuen Chen, Felix Leeb, Van Q. Truong, Punya Syon Pandey, Yves Bicker, Suvajit Majumder, Wenyuan Jiang, Zeju Qiu, Sankalan Pal Chowdhury, Mrinmaya Sachan, Bernhard Schölkopf, Mona T. Diab, Zhijing Jin
| Challenge: | Emerging AI-powered writing assistants focus on grammar fixes or simulating peer review with final scores, yet they fall short of providing concrete, actionable suggestions that help students improve their papers during drafting. |
| Approach: | They propose a human-centered writing assistant system that delivers actionable suggestions as Overleaf-native inline comments while leaving the actual writing entirely to human authors. |
| Outcome: | The proposed system outperforms a baseline with the skill library and provides actionable suggestions while leaving the actual writing to human authors. |
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens (2026.findings-eacl)
Copied to clipboard
| Challenge: | anthropological accounts of culture often focus on static facts or homogeneous values . large language models are being implemented in translation systems, educational tools and search engines . |
| Approach: | They propose to categorize how benchmarks frame culture such as knowledge, preference, performance, or bias. |
| Outcome: | The proposed framework categorizes how benchmarks frame culture, such as knowledge, preference, performance, or bias. |
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Sparse Autoencoders (SAEs) are a promising tool for disentangling FM representations, but they struggle to capture rare, yet crucial concepts in the data. |
| Approach: | They propose a technique to train Sparse Autoencoders to illuminate elusive dark matter features by focusing on specific subdomains. |
| Outcome: | The proposed method achieves 12.5% better classification accuracy than general-purpose SAEs when applied to remove spurious gender information. |
Toward Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST) (2025.findings-acl)
Copied to clipboard
Jiarui Liu, Iman Ouzzani, Wenkai Li, Lechen Zhang, Tianyue Ou, Houda Bouamor, Zhijing Jin, Mona T. Diab
| Challenge: | Despite advances in machine translation, domain-specific terminology translation remains challenging. |
| Approach: | They propose a large-scale multilingual AI terminology dataset that combines LLMs for extraction with human expertise for translation. |
| Outcome: | The proposed framework combines human translation expertise with LLMs to improve translation accuracy and improve BLEU and COMET scores. |
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for embedding human personality traits into LLMs are limited by realism and validity issues. |
| Approach: | They propose to use a large-scale dataset to embed human personality traits into LLMs . they use supervised fine-tuning and direct preference optimization to train LLM models . |
| Outcome: | The proposed methods outperform prompting on personality assessments and IPIP-NEO, and show higher conscientiousness, agreeableness, lower extraversion, and lower neuroticism on reasoning tasks. |
Sentipolis: Emotion-Aware Agents for Social Simulations (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in reasoning and long-context memory are making large language models (LLMs) appear increasingly human-like, which has led researchers to adopt LLM agents as a substrate for social simulation. |
| Approach: | They propose a framework for emotionally stateful agents that integrates continuous Pleasure-Arousal-Dominance representation, dual-speed emotion dynamics, and emotion–memory coupling. |
| Outcome: | The proposed framework improves emotional grounded behavior, boosting communication, and emotional continuity across thousands of interactions over multiple base models and evaluators. |
CoRAG: Collaborative Retrieval-Augmented Generation (2025.naacl-short)
Copied to clipboard
| Challenge: | Existing research on Retrieval-Augmented Generation models has focused on centralized settings where a single entity controls both the model and the datastore. |
| Approach: | They propose a framework for RAG where clients jointly train a shared model using a collaborative passage store. |
| Outcome: | The proposed framework outperforms parametric learning methods and locally trained models in low-resource scenarios. |
Taming Object Hallucinations with Verified Atomic Confidence Estimation (2026.eacl-long)
Copied to clipboard
| Challenge: | Multimodal Large Language Models suffer from hallucinations, especially errors in object existence, attributes, or relations. |
| Approach: | They propose a framework that decomposes responses into atomic queries and estimates confidence using self-consistency or self-confidence aggregation. |
| Outcome: | Experiments on five benchmarks show that TACO outperforms direct prompting and Visual Contrastive Decoding and improves confidence calibration. |
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Language models fail to selectively refuse to answer based on flawed context, study finds . current benchmarks fail to evaluate complex capabilities like selective refusal . |
| Approach: | They propose a framework that generates diagnostic test cases through controlled linguistic perturbation. |
| Outcome: | The proposed framework employs 176 perturbation strategies across six categories of uncertainty and three intensity levels. |
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics (2025.emnlp-main)
Copied to clipboard
Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap
| Challenge: | a study of multi-dimensional persona effects in AI-AI debates shows that personas influence moral stances and debate outcomes . political ideology and personality traits exert the strongest influence, according to our study . |
| Approach: | They propose to use a 6-dimensional persona space to simulate structured debates . they find political ideology and personality traits exert the strongest influence . |
| Outcome: | The study shows that personas affect moral stances and debate outcomes . political ideology and personality traits exert the strongest influence . |
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent work has found that vision-language models trained under the Contrastive Language Image Pre-training framework contain intrinsic social biases, but how these biase relates to downstream performance has been unclear. |
| Approach: | They present the largest comprehensive analysis to-date of how upstream pre-training factors and downstream performance of CLIP models relate to their intrinsic biases. |
| Outcome: | The proposed model performance analysis shows that the choice of pre-training dataset is the most significant upstream predictor of bias, whereas architectural variations have minimal impact. |
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing encoder-based vision-language models (VLMs) contain intrinsic biases that manifest in biased outputs. |
| Approach: | They propose a framework to measure intrinsic bias propagation by correlating intrinsic bias with extrinsic bias in zero-shot text-to-image and image-totext retrieval. |
| Outcome: | The proposed framework shows that larger/better-performing models exhibit greater bias propagation, raising concerns given the trend towards increasingly complex AI models. |
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Modern language models are evaluated on large benchmarks, which are difficult to make sense of. |
| Approach: | They propose a framework to Simplify Benchmark Analysis using model-centric evaluation numbers. |
| Outcome: | The proposed framework can be applied to HELM, MMLU, and BigBenchLite benchmarks. |
From Complexity to Clarity: AI/NLP’s Role in Regulatory Compliance (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in natural language processing have demonstrated remarkable capabilities in text analysis and reasoning. |
| Approach: | They propose to use standardized evaluation frameworks and balanced human-AI collaboration to address these challenges. |
| Outcome: | The proposed research will focus on standardized evaluation frameworks and balanced human-AI collaboration to address these challenges. |
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit anthropomorphism characteristics – human-like qualities portrayed across their outlook, language, behavior, and reasoning functions. |
| Approach: | They propose that anthropomorphism should be treated as a design concept that can be intentionally tuned to support user goals. |
| Outcome: | The proposed design should reflect interaction between artifact designers and interpreters, and should be based on cues embedded in the artifactor and the (cognitive) responses of interpreters to the cue. |