| Challenge: | Existing metrics for conditional language models are not universally comparable . a number of assumptions are made about their quality in multilingual settings . |
| Approach: | They make assumptions about the quality of conditional language models and their semantic meanings that are not consistent with current metrics. |
| Outcome: | The results show that current evaluation metrics are not universally comparable . they use six metrics on two multi-parallel corpora with mono- and multilingual models . |
Similar Papers
On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations (2022.acl-short)
Copied to clipboard
Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, Aram Galstyan
| Challenge: | Recent natural language processing systems use large language models as the backbone . however, societal biases are encoded in these models and transferred to downstream applications . |
| Approach: | They propose to use two categories to measure fairness in natural language processing tasks . they find intrinsic and extrinsic metrics do not correlate in their original setting . |
| Outcome: | The proposed metrics do not correlate in their original setting, the authors show . they find that they are not accurate when correcting for metric misalignments and noise . |
Accounting for Language Effect in the Evaluation of Cross-lingual AMR Parsers (2022.coling-1)
Copied to clipboard
| Challenge: | Existing multilingual AMR evaluation metrics are not available for cross-lingual parsers . existing studies show that source language has a dramatic effect on cross-linguistic AMRs . |
| Approach: | They propose to use three multilingual adaptations of monolingual AMR evaluation metrics to evaluate cross-lingual AML parsers. |
| Outcome: | The proposed metric is the most highly correlated to english AMRs, while the most correlated is S2match. |
A Tour of Explicit Multilingual Semantics: Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing (2022.aacl-tutorials)
Copied to clipboard
| Challenge: | a recent advent of pretrained language models has sparked a revolution in NLP . but, there are still questions about whether current approaches capture explicit, symbolic meaning . this tutorial will review efforts to tackle three key open problems in lexical and sentence-level semantics . |
| Approach: | This tutorial reviews recent efforts to shed light on meaning in NLP . it will focus on three key open problems in lexical and sentence-level semantics . |
| Outcome: | This tutorial reviews recent efforts to shed light on meaning in NLP . it focuses on three key open problems in lexical and sentence-level semantics . |
Multi-lingual Functional Evaluation for Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Multilingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. |
| Approach: | They extend existing functional benchmark templates from English to five additional languages that span the range of resources available for NLP: French, Spanish, Hindi, Arabic and Yoruba. |
| Outcome: | The proposed models are translated from English to French, Spanish, Hindi, Arabic and Yoruba. |
Quantifying Language Disparities in Multilingual Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Contemporary NLP development relies on digital language datasets to build large language models. |
| Approach: | They propose a framework that disentangles confounding variables and introduces interpretable metrics to quantify model performance and language disparities. |
| Outcome: | The proposed framework provides a more reliable measurement of model performance and language disparities for low-resource languages. |
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated. |
| Approach: | They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context. |
| Outcome: | The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks. |
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data (2020.acl-main)
Copied to clipboard
| Challenge: | a priori, large neural language models are described as understanding or capturing meaning on tasks that are ostensibly meaningsensitive. |
| Approach: | They argue that a system trained only on form has no way to learn meaning . they argue that this is due to a misunderstanding of the relationship between form and meaning - which is a misconception in NLP . |
| Outcome: | The proposed model can't learn meaning because it only uses form as training data, the authors argue . they argue that a clear understanding of the distinction between form and meaning will guide the field towards better science around natural language understanding. |
What Meaning-Form Correlation Has to Compose With: A Study of MFC on Artificial and Natural Language (2020.coling-main)
Copied to clipboard
| Challenge: | Compositionality is a widely discussed property of natural languages, although its exact definition has been elusive. |
| Approach: | They propose that compositionality can be measured by measuring meaning-form correlation . they analyze three sets of languages: artificial toy languages tailored to be compositional . |
| Outcome: | The proposed method can assess compositionality on three sets of languages . linguistic phenomena such as synonymy and ungrounded stop-words weigh on the results . |
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist (2023.acl-long)
Copied to clipboard
| Challenge: | a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human . |
| Approach: | They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation . |
| Outcome: | The proposed framework provides access to the evaluation tools for three NLG tasks. |