| Challenge: | Currently, visualizations of worldwide linguistic diversity are limited by point symbology . |
| Approach: | They propose a method to visualize linguistic diversity using point symbology instead of Mercator . instead of languages-as-points, they use Voronoi/Thiessen tessellations to model linguistic areas . |
| Outcome: | The proposed method is based on an Eckert IV projection instead of Mercator . instead of languages-as-points, it uses Voronoi/Thiessen tessellations to model linguistic areas . |
Similar Papers
The State and Fate of Linguistic Diversity and Inclusion in the NLP World (2020.acl-main)
Copied to clipboard
| Challenge: | a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications. |
| Approach: | They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time. |
| Outcome: | The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted. |
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages . |
| Approach: | They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity. |
| Outcome: | The proposed measure can be used to identify the types of languages that are not represented in a data set. |
Mapping 1,000+ Language Models via the Log-Likelihood Vector (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods to compare autoregressive language models are based on log-likelihoods . a model map is constructed using coordinates that capture the geometric structure of probability distributions based upon text-generation probabilities. |
| Approach: | They propose to use log-likelihood vectors to compare autoregressive language models . when treated as model features, their squared Euclidean distance approximates KL divergence . |
| Outcome: | The proposed method is highly scalable and easy to implement. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies . |
| Approach: | They will examine how different socio-cultural perspectives influence what is taken as ground truth by models. |
| Outcome: | This tutorial examines how different socio-cultural perspectives influence representations of global concepts. |
What is ”Typological Diversity” in NLP? (2024.emnlp-main)
Copied to clipboard
| Challenge: | linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims. |
| Approach: | They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages. |
| Outcome: | The proposed measures show that skewed language selection can lead to overestimated multilingual performance. |
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)
Copied to clipboard
| Challenge: | linguistic typology is the classification of languages according to their linguistic properties. |
| Approach: | They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale. |
| Outcome: | The proposed model can predict typological properties on a massively multilingual scale. |
Modeling language evolution and feature dynamics in a realistic geographic environment (2020.coling-main)
Copied to clipboard
| Challenge: | a number of studies have examined the stability or biases of typological features within language families . |
| Approach: | They propose a model for simulating languages and their features over time in a realistic geographic environment. |
| Outcome: | The proposed model is flexible and realistic, and can be used to answer questions. |
The Geometry of Multilingual Language Model Representations (2022.emnlp-main)
Copied to clipboard
| Challenge: | XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning. |
| Approach: | They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language. |
| Outcome: | The proposed model can extract features for downstream tasks and cross-lingual transfer learning. |
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Language models (LMs) have exhibited impressive abilities in generating code from natural language requirements. |
| Approach: | They propose to introduce various metrics with inter-code similarity to evaluate the diversity of generated code by comparing model-generated solutions with human-written ones. |
| Outcome: | The proposed method leverages LMs’ capabilities in code understanding and reasoning, resulting in a set of metrics that represent the number of algorithms in model-generated solutions. |