| Challenge: | a new method is being developed to assess how well each language is doing in terms of digital language support. |
| Approach: | They develop an automated method to assess how well each language is doing in terms of digital language support. |
| Outcome: | The proposed method scrapes the names of supported languages from 143 digital tools and produces an explainable model for quantifying and monitoring it on a global scale. |
Similar Papers
Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)
Copied to clipboard
| Challenge: | Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages. |
| Approach: | They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
| Outcome: | The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
High Performance Natural Language Processing (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | a tutorial on scaling natural language processing will recapitulate the state-of-the-art in the field . |
| Approach: | This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective. |
| Outcome: | This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective. |
Global Readiness of Language Technology for Healthcare: What Would It Take to Combat the Next Pandemic? (2022.coling-1)
Copied to clipboard
| Challenge: | Language Technology (LT) has been used in the COVID-19 pandemic, but only in a handful of languages. |
| Approach: | They propose to use conversational agents for information dissemination and basic diagnosis in 15 Asian and African languages with varying resource-availability to test their knowledge of LT. |
| Outcome: | The proposed research confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help prioritize research and investment strategies in LT for healthcare. |
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)
Copied to clipboard
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang
| Challenge: | Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains. |
| Approach: | They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks . |
| Outcome: | The proposed evaluations are reproducible, reliable, and robust. |
Language Technology for Multilingual Europe: An Analysis of a Large-Scale Survey regarding Challenges, Demands, Gaps and Needs (L18-1)
Copied to clipboard
| Challenge: | a survey titled "Language Technology for Multilingual Europe" was conducted between May and June 2017 . 634 participants in 52 countries responded to the survey . |
| Approach: | a large-scale survey was conducted to assess the best multilingual technologies in Europe. a total of 634 participants in 52 countries responded to the survey. |
| Outcome: | The study aims to identify the biggest challenges, obstacles and gaps in European language technology . participants were encouraged to share concrete suggestions and recommendations on how present challenges can be turned into opportunities . |
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)
Copied to clipboard
| Challenge: | Large text corpora are increasingly important for a wide variety of NLP tasks. |
| Approach: | They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages. |
| Outcome: | The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation. |
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)
Copied to clipboard
Daan van Esch, Sandy Ritchie, Sebastian Ruder, Julia Kreutzer, Clara Rivera, Ishank Saxena, Isaac Caswell
| Challenge: | Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages . |
| Approach: | They provide an overview of some of the major online data sources available for thousands of languages. |
| Outcome: | The proposed language technologies are based on the data available for thousands of languages. |
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)
Copied to clipboard
| Challenge: | introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance. |
| Approach: | They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them. |
| Outcome: | The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods. |
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment. |
| Approach: | They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge. |
| Outcome: | The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |