Form and Meaning in Intrinsic Multilingual Evaluations (2026.eacl-long)

Copied to clipboard

Challenge: Existing metrics for conditional language models are not universally comparable . a number of assumptions are made about their quality in multilingual settings .
Approach: They make assumptions about the quality of conditional language models and their semantic meanings that are not consistent with current metrics.
Outcome: The results show that current evaluation metrics are not universally comparable . they use six metrics on two multi-parallel corpora with mono- and multilingual models .

Similar Papers

On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations (2022.acl-short)

Copied to clipboard

Challenge: Recent natural language processing systems use large language models as the backbone . however, societal biases are encoded in these models and transferred to downstream applications .
Approach: They propose to use two categories to measure fairness in natural language processing tasks . they find intrinsic and extrinsic metrics do not correlate in their original setting .
Outcome: The proposed metrics do not correlate in their original setting, the authors show . they find that they are not accurate when correcting for metric misalignments and noise .
Accounting for Language Effect in the Evaluation of Cross-lingual AMR Parsers (2022.coling-1)

Copied to clipboard

Challenge: Existing multilingual AMR evaluation metrics are not available for cross-lingual parsers . existing studies show that source language has a dramatic effect on cross-linguistic AMRs .
Approach: They propose to use three multilingual adaptations of monolingual AMR evaluation metrics to evaluate cross-lingual AML parsers.
Outcome: The proposed metric is the most highly correlated to english AMRs, while the most correlated is S2match.
A Tour of Explicit Multilingual Semantics: Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing (2022.aacl-tutorials)

Copied to clipboard

Challenge: a recent advent of pretrained language models has sparked a revolution in NLP . but, there are still questions about whether current approaches capture explicit, symbolic meaning . this tutorial will review efforts to tackle three key open problems in lexical and sentence-level semantics .
Approach: This tutorial reviews recent efforts to shed light on meaning in NLP . it will focus on three key open problems in lexical and sentence-level semantics .
Outcome: This tutorial reviews recent efforts to shed light on meaning in NLP . it focuses on three key open problems in lexical and sentence-level semantics .
Multi-lingual Functional Evaluation for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Multilingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM.
Approach: They extend existing functional benchmark templates from English to five additional languages that span the range of resources available for NLP: French, Spanish, Hindi, Arabic and Yoruba.
Outcome: The proposed models are translated from English to French, Spanish, Hindi, Arabic and Yoruba.
Quantifying Language Disparities in Multilingual Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Contemporary NLP development relies on digital language datasets to build large language models.
Approach: They propose a framework that disentangles confounding variables and introduces interpretable metrics to quantify model performance and language disparities.
Outcome: The proposed framework provides a more reliable measurement of model performance and language disparities for low-resource languages.
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated.
Approach: They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context.
Outcome: The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks.
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data (2020.acl-main)

Copied to clipboard

Challenge: a priori, large neural language models are described as understanding or capturing meaning on tasks that are ostensibly meaningsensitive.
Approach: They argue that a system trained only on form has no way to learn meaning . they argue that this is due to a misunderstanding of the relationship between form and meaning - which is a misconception in NLP .
Outcome: The proposed model can't learn meaning because it only uses form as training data, the authors argue . they argue that a clear understanding of the distinction between form and meaning will guide the field towards better science around natural language understanding.
What Meaning-Form Correlation Has to Compose With: A Study of MFC on Artificial and Natural Language (2020.coling-main)

Copied to clipboard

Challenge: Compositionality is a widely discussed property of natural languages, although its exact definition has been elusive.
Approach: They propose that compositionality can be measured by measuring meaning-form correlation . they analyze three sets of languages: artificial toy languages tailored to be compositional .
Outcome: The proposed method can assess compositionality on three sets of languages . linguistic phenomena such as synonymy and ungrounded stop-words weigh on the results .
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist (2023.acl-long)

Copied to clipboard

Challenge: a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human .
Approach: They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation .
Outcome: The proposed framework provides access to the evaluation tools for three NLG tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations