Towards faithfully visualizing global linguistic diversity (L18-1)

Copied to clipboard

Challenge: Currently, visualizations of worldwide linguistic diversity are limited by point symbology .
Approach: They propose a method to visualize linguistic diversity using point symbology instead of Mercator . instead of languages-as-points, they use Voronoi/Thiessen tessellations to model linguistic areas .
Outcome: The proposed method is based on an Eckert IV projection instead of Mercator . instead of languages-as-points, it uses Voronoi/Thiessen tessellations to model linguistic areas .

Similar Papers

The State and Fate of Linguistic Diversity and Inclusion in the NLP World (2020.acl-main)

Copied to clipboard

Challenge: a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications.
Approach: They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time.
Outcome: The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted.
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)

Copied to clipboard

Challenge: a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages .
Approach: They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity.
Outcome: The proposed measure can be used to identify the types of languages that are not represented in a data set.
Mapping 1,000+ Language Models via the Log-Likelihood Vector (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to compare autoregressive language models are based on log-likelihoods . a model map is constructed using coordinates that capture the geometric structure of probability distributions based upon text-generation probabilities.
Approach: They propose to use log-likelihood vectors to compare autoregressive language models . when treated as model features, their squared Euclidean distance approximates KL divergence .
Outcome: The proposed method is highly scalable and easy to implement.
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)

Copied to clipboard

Challenge: Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages.
Approach: They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices .
Outcome: The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices.
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)

Copied to clipboard

Challenge: audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies .
Approach: They will examine how different socio-cultural perspectives influence what is taken as ground truth by models.
Outcome: This tutorial examines how different socio-cultural perspectives influence representations of global concepts.
What is ”Typological Diversity” in NLP? (2024.emnlp-main)

Copied to clipboard

Challenge: linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims.
Approach: They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages.
Outcome: The proposed measures show that skewed language selection can lead to overestimated multilingual performance.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
Modeling language evolution and feature dynamics in a realistic geographic environment (2020.coling-main)

Copied to clipboard

Challenge: a number of studies have examined the stability or biases of typological features within language families .
Approach: They propose a model for simulating languages and their features over time in a realistic geographic environment.
Outcome: The proposed model is flexible and realistic, and can be used to answer questions.
The Geometry of Multilingual Language Model Representations (2022.emnlp-main)

Copied to clipboard

Challenge: XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning.
Approach: They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language.
Outcome: The proposed model can extract features for downstream tasks and cross-lingual transfer learning.
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code (2025.findings-emnlp)

Copied to clipboard

Challenge: Language models (LMs) have exhibited impressive abilities in generating code from natural language requirements.
Approach: They propose to introduce various metrics with inter-code similarity to evaluate the diversity of generated code by comparing model-generated solutions with human-written ones.
Outcome: The proposed method leverages LMs’ capabilities in code understanding and reasoning, resulting in a set of metrics that represent the number of algorithms in model-generated solutions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations