Challenge: Recent work in linguistics and NLP has investigated the quantity and quality of AAL representation in pretraining corpora.
Approach: They examine the quantity and quality of African American Language (AAL) representation in pretraining corpora.
Outcome: The results show that AAL is underrepresented in all evaluated corpora compared to US demographics . they also show that most automated filters are more likely to conserve white Mainstream English (WME) texts over AAL .

Similar Papers

My LLM might Mimic AAE - But When Should It? (2025.naacl-long)

Copied to clipboard

Challenge: a study examines the representation of African American English in large language models . a survey of black americans and annotation of LLM outputs shows that Black Americans prefer to use AAE in formal settings .
Approach: They examine Black Americans' perceptions of how effective AI tools are at producing authentic African American English in large language models.
Outcome: The results show that Black Americans prefer to use LLMs in formal settings over informal ones . the results show they prefer to produce AAE in less formal settings .
Evaluation of African American Language Bias in Natural Language Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages.
Approach: They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM.
Outcome: The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model.
Better Quality Pre-training Data and T5 Models for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing web crawls have demonstrated quality issues for low-resource languages . Existing pretraining corpora have numerous quality issues .
Approach: They propose to audit existing pretraining corpora to understand and rectify quality issues . they pretrain a new T5-based model and evaluate its performance on multiple tasks .
Outcome: The proposed model outperforms existing pretrained models on four NLP tasks.
Rejected Dialects: Biases Against African American Language in Reward Models (2025.findings-naacl)

Copied to clipboard

Challenge: Preference alignment via reward models can introduce new biases, hindering reward models’ fairness and equity.
Approach: They propose a framework for evaluating dialect biases in reward models and conduct a case study on biase . they compare reward models' preferences and behavior on paired White Mainstream English and machine-translated and human-written AAL corpora.
Outcome: The proposed framework evaluates dialect biases in reward models and compares them with paired White Mainstream English (WME) and machine-translated and human-written AAL corpora.
Investigating African-American Vernacular English in Transformer-Based Text Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work in Natural Language Generation (NLG) uses a Transformer-based language model to generate high-quality, coherent text when prompted by arbitrary input.
Approach: They evaluate the performance of a Transformer-based model that generates high-quality, coherent text when prompted by arbitrary input.
Outcome: The proposed model improves on AAVE and SAE text with pretrained sentiment classifiers.
Leveraging Syntactic Dependencies in Disambiguation: The Case of African American English (2024.lrec-main)

Copied to clipboard

Challenge: African American English (AAE) is a low-resource language facing the challenge of inadequate annotated data for training natural language processing models.
Approach: They propose a syntactically informed classifier for automatic disambiguation of AAE's habitual be.
Outcome: The proposed classifier improves automatic disambiguation of habitual and non-habitual meanings of "be" integrating syntactic information improves disambiguations of habituality by 65 F1 points over baseline models and as much as 74 points.
Evaluating and Mitigating Inherent Linguistic Bias of African American English through Inference (2022.coling-1)

Copied to clipboard

Challenge: Recent studies show that NLP models trained on standard English produce biased outcomes against underrepresented English varieties.
Approach: They propose a morphosyntactically-informed rule-based translation method that uses a greedy algorithm to debiase NLP models.
Outcome: The proposed framework outperforms large language models while maintaining or improving the prediction performance.
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies have suggested that the composition of the pretraining corpus exerts a significant impact upon the performance of LLMs.
Approach: They analyze the impact of 48 datasets from 5 major categories of pretraining data of Large Language Models and measure their impacts on LLMs using benchmarks about nine major categories.
Outcome: The proposed analysis provides insights into the organization of data to support more efficient pretraining of Large Language Models.
Understanding the Impacts of Language Technologies’ Performance Disparities on African American Language Speakers (2024.findings-acl)

Copied to clipboard

Challenge: Previous work has examined performance disparities between AAL speakers and White Mainstream English speakers . but, this work has not sought to understand the impacts of these disparities on AAL speaker.
Approach: They examine the experiences of African American Language (AAL) speakers when using language technologies.
Outcome: The authors interview 19 AAL speakers to understand performance disparities . they find that speakers often undertake invisible labor to successfully use language technologies .
Analysis of LLM as a grammatical feature tagger for African American English (2025.findings-naacl)

Copied to clipboard

Challenge: African American English (AAE) presents unique challenges in natural language processing (NLP).
Approach: They evaluate the ability of different NLP systems to recognize distinctive AAE grammatical features by using sentence-level binary classification tasks using both zero-shot and fewshot strategies.
Outcome: The evaluation involved sentence-level binary classification tasks, using both zero-shot and few-shot strategies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations