URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base (2025.coling-main)
Copied to clipboard
Aditya Armaan Khan, Mason Stephen Shipton, David Anugraha, Kaiyao Duan, Phuong H. Hoang, Eric Khiu, A. Seza Doğruöz, Annie Lee
| Challenge: | URIEL is limited in terms of linguistic inclusion and overall usability . URIel+ provides robust, customizable distance calculations to better suit the needs of users. |
| Approach: | They propose a new version of URIEL and a query tool that provides a standardized approach to representing languages as geographical, phylogenetic, and typological vectors. |
| Outcome: | URIEL+ expands the user experience with robust, customizable distance calculations to better suit the needs of users. |
Similar Papers
Less is More: The Effectiveness of Compact Typological Language Representations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Linguistic feature datasets such as URIEL+ have high dimensionality and sparsity, especially for low-resource languages. |
| Approach: | They propose a pipeline to optimize the URIEL+ typological feature space by feature selection and imputation. |
| Outcome: | The proposed pipeline produces compact yet interpretable typological representations on linguistic distance alignment and downstream tasks. |
Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+ (2026.eacl-srw)
Copied to clipboard
York Hay Ng, Aditya Khan, Xiang Lu, Matteo Salloum, Michael Zhou, Phuong Hanh Hoang, A. Seza Doğruöz, En-Shiun Annie Lee
| Challenge: | Existing linguistic knowledge bases such as URIEL+ lack a principled method for aggregating these signals into a single, comprehensive score. |
| Approach: | They propose a framework for type-matched language distances that unifies these signals into a robust, task-agnostic composite distance. |
| Outcome: | The proposed representations improve transfer performance when the distance type is relevant to the task, while yielding gains in most tasks. |
data2lang2vec: Data Driven Typological Features Completion (2025.coling-main)
Copied to clipboard
| Challenge: | Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database. |
| Approach: | They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features. |
| Outcome: | The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms. |
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | . - (EN) |
| Approach: | . - (EN) |
| Outcome: | . - (EN) |
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)
Copied to clipboard
| Challenge: | Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing . |
| Approach: | They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser. |
| Outcome: | The proposed model can be trained on 40 languages with the help of typological features. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Findings of the Association for Computational Linguistics: NAACL 2022 (2022.findings-naacl)
Copied to clipboard
| Challenge: | . - (EN) |
| Approach: | . - (EN) |
| Outcome: | . - (EN) |
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 5: Tutorial Abstracts) (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | NAACL 2024 tutorial sessions are a cornerstone event of the conference . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
| Approach: | NAACL 2024 will host a tutorial session featuring top-notch researchers . the tutorials aim to equip attendees with the latest tools and methodologies . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
| Outcome: | the tutorial sessions are a cornerstone event of the conference . the call, submission, reviewing, and selection of tutorials were coordinated . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |