| Challenge: | a study examines the representation of African American English in large language models . a survey of black americans and annotation of LLM outputs shows that Black Americans prefer to use AAE in formal settings . |
| Approach: | They examine Black Americans' perceptions of how effective AI tools are at producing authentic African American English in large language models. |
| Outcome: | The results show that Black Americans prefer to use LLMs in formal settings over informal ones . the results show they prefer to produce AAE in less formal settings . |
Similar Papers
Evaluation of African American Language Bias in Natural Language Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages. |
| Approach: | They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM. |
| Outcome: | The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model. |
Data Caricatures: On the Representation of African American Language in Pretraining Corpora (2025.acl-long)
Copied to clipboard
Nicholas Deas, Blake Vente, Amith Ananthram, Jessica A Grieser, Desmond U. Patton, Shana Kleiner, James R. Shepard Iii, Kathleen McKeown
| Challenge: | Recent work in linguistics and NLP has investigated the quantity and quality of AAL representation in pretraining corpora. |
| Approach: | They examine the quantity and quality of African American Language (AAL) representation in pretraining corpora. |
| Outcome: | The results show that AAL is underrepresented in all evaluated corpora compared to US demographics . they also show that most automated filters are more likely to conserve white Mainstream English (WME) texts over AAL . |
Analysis of LLM as a grammatical feature tagger for African American English (2025.findings-naacl)
Copied to clipboard
| Challenge: | African American English (AAE) presents unique challenges in natural language processing (NLP). |
| Approach: | They evaluate the ability of different NLP systems to recognize distinctive AAE grammatical features by using sentence-level binary classification tasks using both zero-shot and fewshot strategies. |
| Outcome: | The evaluation involved sentence-level binary classification tasks, using both zero-shot and few-shot strategies. |
Leveraging Syntactic Dependencies in Disambiguation: The Case of African American English (2024.lrec-main)
Copied to clipboard
| Challenge: | African American English (AAE) is a low-resource language facing the challenge of inadequate annotated data for training natural language processing models. |
| Approach: | They propose a syntactically informed classifier for automatic disambiguation of AAE's habitual be. |
| Outcome: | The proposed classifier improves automatic disambiguation of habitual and non-habitual meanings of "be" integrating syntactic information improves disambiguations of habituality by 65 F1 points over baseline models and as much as 74 points. |
Rejected Dialects: Biases Against African American Language in Reward Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Preference alignment via reward models can introduce new biases, hindering reward models’ fairness and equity. |
| Approach: | They propose a framework for evaluating dialect biases in reward models and conduct a case study on biase . they compare reward models' preferences and behavior on paired White Mainstream English and machine-translated and human-written AAL corpora. |
| Outcome: | The proposed framework evaluates dialect biases in reward models and compares them with paired White Mainstream English (WME) and machine-translated and human-written AAL corpora. |
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated. |
| Approach: | They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context. |
| Outcome: | The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks. |
Evaluating and Mitigating Inherent Linguistic Bias of African American English through Inference (2022.coling-1)
Copied to clipboard
| Challenge: | Recent studies show that NLP models trained on standard English produce biased outcomes against underrepresented English varieties. |
| Approach: | They propose a morphosyntactically-informed rule-based translation method that uses a greedy algorithm to debiase NLP models. |
| Outcome: | The proposed framework outperforms large language models while maintaining or improving the prediction performance. |
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories. |
| Approach: | 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes . |
| Outcome: | 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes . |
Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | Language ideologies are evaluative ideas or beliefs about language, such as ideas about what is "correct", "natural" or "articulate". |
| Approach: | They use gender-neutral variants more often when more explicit metalinguistic context is provided. |
| Outcome: | The findings show that language ideologies in LLMs can vary, which may be unexpected to users. |
It’s Not Bragging If You Can Back It Up: Can LLMs Understand Braggings? (2025.acl-long)
Copied to clipboard
| Challenge: | Bragging is a pervasive social-linguistic phenomenon that reflects complex human interaction patterns. |
| Approach: | They propose to use bragging recognition, bragging explanation, and bragging generation tasks to examine bragging in large language models (LLMs) . |
| Outcome: | The proposed models can identify bragging intent, social appropriateness, and account for context sensitivity and provide new insights into how LLMs process bragging. |