Papers with AAVE
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)
Copied to clipboard
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei
| Challenge: | a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination . |
| Approach: | They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE . |
| Outcome: | The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE. |
Investigating African-American Vernacular English in Transformer-Based Text Generation (2020.emnlp-main)
Copied to clipboard
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, William Yang Wang
| Challenge: | Recent work in Natural Language Generation (NLG) uses a Transformer-based language model to generate high-quality, coherent text when prompted by arbitrary input. |
| Approach: | They evaluate the performance of a Transformer-based model that generates high-quality, coherent text when prompted by arbitrary input. |
| Outcome: | The proposed model improves on AAVE and SAE text with pretrained sentiment classifiers. |
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations (2026.acl-long)
Copied to clipboard
| Challenge: | Agentic benchmarks rely on LLM-simulated users to evaluate agent performance . however, the robustness, validity, and fairness of this approach remain unexamined . |
| Approach: | They investigate whether LLM-simulated users are reliable proxies for real human users . they find that agent success rates vary up to 9 percentage points across different LLMs . |
| Outcome: | The results show that simulated users underestimate success on challenging tasks while miscalibrate performance on moderately difficult tasks. |