| Challenge: | Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database. |
| Approach: | They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features. |
| Outcome: | The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms. |
Similar Papers
URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base (2025.coling-main)
Copied to clipboard
Aditya Armaan Khan, Mason Stephen Shipton, David Anugraha, Kaiyao Duan, Phuong H. Hoang, Eric Khiu, A. Seza Doğruöz, Annie Lee
| Challenge: | URIEL is limited in terms of linguistic inclusion and overall usability . URIel+ provides robust, customizable distance calculations to better suit the needs of users. |
| Approach: | They propose a new version of URIEL and a query tool that provides a standardized approach to representing languages as geographical, phylogenetic, and typological vectors. |
| Outcome: | URIEL+ expands the user experience with robust, customizable distance calculations to better suit the needs of users. |
Multilingual Gradient Word-Order Typology from Universal Dependencies (2024.eacl-short)
Copied to clipboard
| Challenge: | Existing typological databases, including WALS and Grambank, suffer from inconsistencies due to categorical format. |
| Approach: | They propose a new seed dataset that uses continuous-valued data instead of categorical data to better reflect the variability of language. |
| Outcome: | The proposed dataset can be easily adapted to generate data for a broader set of features and languages. |
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing evidence on the intrinsic difficulty of multilingual modeling is limited to small monolingual models or bilingual models trained from scratch. |
| Approach: | They propose to use typological properties to determine the difficulty of modeling a language . they analyze two large pre-trained multilingual translation models . |
| Outcome: | The proposed models are based on two large pre-trained models of encoder-decoder and decoder-only machine translation. |
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)
Copied to clipboard
| Challenge: | Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing . |
| Approach: | They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser. |
| Outcome: | The proposed model can be trained on 40 languages with the help of typological features. |
Less is More: The Effectiveness of Compact Typological Language Representations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Linguistic feature datasets such as URIEL+ have high dimensionality and sparsity, especially for low-resource languages. |
| Approach: | They propose a pipeline to optimize the URIEL+ typological feature space by feature selection and imputation. |
| Outcome: | The proposed pipeline produces compact yet interpretable typological representations on linguistic distance alignment and downstream tasks. |
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars. |
| Approach: | They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources. |
| Outcome: | The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios. |
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning. |
| Approach: | They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization. |
| Outcome: | The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors. |
Working Hard or Hardly Working: Challenges of Integrating Typology into Neural Dependency Parsers (D19-1)
Copied to clipboard
| Challenge: | linguistic typology has shown great promise in pre-neural parsing, but results for neural architectures have been mixed. |
| Approach: | They explore the task of leveraging typology in the context of cross-lingual dependency parsing. |
| Outcome: | The proposed approach improves performance in the context of cross-lingual dependency parsing. |
Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies consider linguistic typology as a potential source of knowledge to support multilingual natural language processing (NLP) tasks. |
| Approach: | They propose to fuse both views using canonical correlation analysis and use it to infer typological features and language phylogenies to construct a multi-view language vector space for multilingual machine translation. |
| Outcome: | The proposed model achieves competitive translation accuracy in multilingual machine translation tasks without expensive retraining of massive multilingual or ranking models. |
A Probabilistic Generative Model of Linguistic Typology (N19-1)
Copied to clipboard
| Challenge: | a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages. |
| Approach: | They propose a generative model of language based on exponential-family matrix factorisation. |
| Outcome: | a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features. |