Assessing Multilinguality of Publicly Accessible Websites (2022.lrec-1)

Copied to clipboard

Challenge: multilingualism on the Web is a problem not only at the world level, but also at the European and regional level.
Approach: They propose a tool that automatically analyses the language diversity of the Web and propose indicators and methodologies to measure multilingualism of European websites.
Outcome: The proposed tool can be independently run at set intervals and concludes that multilingualism on the Web is still a problem not only at the world level, but also at the European and regional level.

Similar Papers

Language Technology for Multilingual Europe: An Analysis of a Large-Scale Survey regarding Challenges, Demands, Gaps and Needs (L18-1)

Copied to clipboard

Challenge: a survey titled "Language Technology for Multilingual Europe" was conducted between May and June 2017 . 634 participants in 52 countries responded to the survey .
Approach: a large-scale survey was conducted to assess the best multilingual technologies in Europe. a total of 634 participants in 52 countries responded to the survey.
Outcome: The study aims to identify the biggest challenges, obstacles and gaps in European language technology . participants were encouraged to share concrete suggestions and recommendations on how present challenges can be turned into opportunities .
Assessing Digital Language Support on a Global Scale (2022.coling-1)

Copied to clipboard

Challenge: a new method is being developed to assess how well each language is doing in terms of digital language support.
Approach: They develop an automated method to assess how well each language is doing in terms of digital language support.
Outcome: The proposed method scrapes the names of supported languages from 143 digital tools and produces an explainable model for quantifying and monitoring it on a global scale.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
An Open Multilingual System for Scoring Readability of Wikipedia (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on the readability of Wikipedia have focused on English only and there are currently no systems supporting automatic readability assessment of the 300+ languages in Wikipedia.
Approach: They propose a multilingual model to assess Wikipedia's readability using a dataset spanning 14 languages.
Outcome: The proposed model outperforms existing models in a zero-shot scenario and is more accurate than previous benchmarks.
ELRC Action: Covering Confidentiality, Correctness and Cross-linguality (2022.lrec-1)

Copied to clipboard

Challenge: ELRC aims to reduce language barriers by assessing language technology (LT) specifications . automated anonymisation and multilingual fake news processing are two of the most extensive LT assessments .
Approach: They describe language technology (LT) assessments carried out by the European Commission . they zoom in on two of the most extensive assessments, namely automated anonymisation and multilingual fake news processing.
Outcome: The language technology (LT) assessments carried out by the European Commission are detailed in this paper . they include a consultation round with stakeholders from public organisations, academia and industry . the ELRC action aims to create proof-of-concept environments integrating relevant tools and services .
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on English and a handful of high-resource languages, thereby overlooking the realistic performance of large language models in multilingual and lower-resourced scenarios.
Approach: They propose a unified and lightweight framework that integrates 27 benchmarks under a standard ISO 639-3 language identifier system to enable seamless incorporation of new benchmarks.
Outcome: The proposed framework integrates 27 benchmarks under a standard ISO 639-3 language identifier system, allowing for seamless incorporation of new benchmarks.
Quantifying Language Disparities in Multilingual Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Contemporary NLP development relies on digital language datasets to build large language models.
Approach: They propose a framework that disentangles confounding variables and introduces interpretable metrics to quantify model performance and language disparities.
Outcome: The proposed framework provides a more reliable measurement of model performance and language disparities for low-resource languages.
The DLDP Survey on Digital Use and Usability of EU Regional and Minority Languages (L18-1)

Copied to clipboard

Challenge: the survey was launched by the Digital Language Diversity Project to investigate the real usage, needs and expectations of European minority language speakers regarding digital opportunities.
Approach: This paper reports on the results of an exploratory survey launched by the Digital Language Diversity Project about the digital use and usability of regional and minority languages on digital media and devices.
Outcome: The findings of the first exploratory survey on the use and usability of regional and minority languages on digital media and devices are presented in this paper.
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated.
Approach: They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context.
Outcome: The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks.
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage.
Approach: They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement.
Outcome: The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations