Challenge: Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved.
Approach: EnDive evaluates seven state-of-the-art large language models across tasks . human evaluations confirm high translation quality, with average scores of at least 6.02/7 .
Outcome: EnDive evaluates state-of-the-art large language models across language understanding, reasoning, mathematics, logic tasks.

Similar Papers

Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)

Copied to clipboard

Challenge: a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination .
Approach: They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE .
Outcome: The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) pre-trained on massive text data in many languages are preferred solution for various Natural Language processing tasks.
Approach: They compare tokenization parity and information parity as representational biases in pre-trained models . they find TP is better predictor of performance on tasks reliant on syntactic and morphological cues .
Outcome: The proposed model improves on dialect classification, topic classification, and extractive question answering tasks.
F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing fairness evaluation benchmarks for large language models rely on closed-ended evaluation formats that overlook factuality considerations rooted in historical, social, physiological, and cultural contexts.
Approach: They propose an open-ended fairness evaluation benchmark for large language models . they incorporate factuality considerations and multi-turn reasoning into the benchmark .
Outcome: The proposed benchmark incorporates factual grounding and text generation to better reflect the complexities of real-world model usage.
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated.
Approach: They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context.
Outcome: The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks.
Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Standard bias benchmarks are used for large language models to measure the association between social attributes and single-word outputs.
Approach: They adapt three standard bias metrics of next-word prediction to measure gender-occupation bias and develop an analogous RUTEd evaluation in three contexts of real-world LLM use.
Outcome: The proposed benchmarks are robust to lengthening model outputs via a more realistic user prompt in the domain of gender-occupation bias.
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains.
Approach: They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs.
Outcome: The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs.
Multi-VALUE: A Framework for Cross-Dialectal English NLP (2023.acl-long)

Copied to clipboard

Challenge: Current systems that focus on standard American English are not dialect invariant . current systems focus on a single dialect, which results in performance discrepancies .
Approach: They propose a resource for evaluating and achieving English dialect invariance . they stress test question answering, machine translation, and semantic parsing .
Outcome: The proposed system is based on a rule-based translation system spanning 50 English dialects and 189 unique linguistic features.
mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluations focus on English-to-Python conversion tasks with limited test cases . code generation from low-resource language prompts remains largely unexplored .
Approach: They propose a benchmark that supports prompts in over 200 natural languages . they provide expert human translations for 15 diverse natural languages (NLs)
Outcome: The HumanEval Benchmark is the most widely used code generation benchmark . it provides expert human translations for 15 diverse natural languages .
FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing (2022.acl-long)

Copied to clipboard

Challenge: Using pre-trained language models, we evaluate performance group disparities while none of these techniques guarantee fairness, nor consistently mitigate group disparity.
Approach: They present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks.
Outcome: The proposed methods show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations