Challenge: Linguistic disparity in the NLP world is widely acknowledged, but the reasons behind it are rarely discussed within the field.
Approach: They propose to categorise languages based on speaker population and vitality . they also analyse the distribution of language data resources and amount of NLP/CL research .
Outcome: The proposed model identifies the reasons for the disparity and suggests ways to overcome it.

Similar Papers

Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages.
Approach: They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Outcome: The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
The State and Fate of Linguistic Diversity and Inclusion in the NLP World (2020.acl-main)

Copied to clipboard

Challenge: a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications.
Approach: They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time.
Outcome: The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted.
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)

Copied to clipboard

Challenge: Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages.
Approach: They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices .
Outcome: The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
Fairness in Language Models Beyond English: Gaps and Challenges (2023.findings-eacl)

Copied to clipboard

Challenge: Language models are inequitable at encoding and re-presentation, but there is much to be studied and criticism for the existing research that remains to be addressed.
Approach: They propose to survey fairness in multilingual and non-English contexts . they argue that it is infeasible to achieve comprehensive coverage in terms of fairness datasets based on English .
Outcome: The proposed methods are infeasible to scale across languages and cultures, the authors argue . they argue that the current methods are too narrowly focused on specific dimensions and types of biases and cannot scale across cultures.
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)

Copied to clipboard

Challenge: There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea.
Approach: They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world.
Outcome: The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Should We Ban English NLP for a Year? (2022.emnlp-main)

Copied to clipboard

Challenge: aaron carroll: two thirds of NLP research is devoted to developing technology for speakers of English . carroll says this bias feeds into consumer technologies to widen existing inequality gaps . he says we need to consider more concrete measures to mitigate climate change .
Approach: a new paper argues that NLP is contributing to global inequalities through a digital language divide . a carbon tax, cap-and-trade and car-free Sundays are examples of measures to mitigate climate change .
Outcome: a new paper argues that NLP is contributing to global inequalities through a digital language divide . a carbon tax, cap-and-trade and car-free Sundays are examples of measures to mitigate climate change .
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)

Copied to clipboard

Challenge: linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages .
Approach: They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions .
Outcome: The proposed model is representative of the language diversity and coverage of natural language processing systems.
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations