data2lang2vec: Data Driven Typological Features Completion (2025.coling-main)

Copied to clipboard

Challenge: Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database.
Approach: They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features.
Outcome: The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms.

Similar Papers

URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base (2025.coling-main)

Copied to clipboard

Challenge: URIEL is limited in terms of linguistic inclusion and overall usability . URIel+ provides robust, customizable distance calculations to better suit the needs of users.
Approach: They propose a new version of URIEL and a query tool that provides a standardized approach to representing languages as geographical, phylogenetic, and typological vectors.
Outcome: URIEL+ expands the user experience with robust, customizable distance calculations to better suit the needs of users.
Multilingual Gradient Word-Order Typology from Universal Dependencies (2024.eacl-short)

Copied to clipboard

Challenge: Existing typological databases, including WALS and Grambank, suffer from inconsistencies due to categorical format.
Approach: They propose a new seed dataset that uses continuous-valued data instead of categorical data to better reflect the variability of language.
Outcome: The proposed dataset can be easily adapted to generate data for a broader set of features and languages.
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing evidence on the intrinsic difficulty of multilingual modeling is limited to small monolingual models or bilingual models trained from scratch.
Approach: They propose to use typological properties to determine the difficulty of modeling a language . they analyze two large pre-trained multilingual translation models .
Outcome: The proposed models are based on two large pre-trained models of encoder-decoder and decoder-only machine translation.
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)

Copied to clipboard

Challenge: Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing .
Approach: They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser.
Outcome: The proposed model can be trained on 40 languages with the help of typological features.
Less is More: The Effectiveness of Compact Typological Language Representations (2025.emnlp-main)

Copied to clipboard

Challenge: Linguistic feature datasets such as URIEL+ have high dimensionality and sparsity, especially for low-resource languages.
Approach: They propose a pipeline to optimize the URIEL+ typological feature space by feature selection and imputation.
Outcome: The proposed pipeline produces compact yet interpretable typological representations on linguistic distance alignment and downstream tasks.
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars.
Approach: They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources.
Outcome: The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios.
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning.
Approach: They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization.
Outcome: The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors.
Working Hard or Hardly Working: Challenges of Integrating Typology into Neural Dependency Parsers (D19-1)

Copied to clipboard

Challenge: linguistic typology has shown great promise in pre-neural parsing, but results for neural architectures have been mixed.
Approach: They explore the task of leveraging typology in the context of cross-lingual dependency parsing.
Outcome: The proposed approach improves performance in the context of cross-lingual dependency parsing.
Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations (2020.emnlp-main)

Copied to clipboard

Challenge: Recent studies consider linguistic typology as a potential source of knowledge to support multilingual natural language processing (NLP) tasks.
Approach: They propose to fuse both views using canonical correlation analysis and use it to infer typological features and language phylogenies to construct a multi-view language vector space for multilingual machine translation.
Outcome: The proposed model achieves competitive translation accuracy in multilingual machine translation tasks without expensive retraining of massive multilingual or ranking models.
A Probabilistic Generative Model of Linguistic Typology (N19-1)

Copied to clipboard

Challenge: a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages.
Approach: They propose a generative model of language based on exponential-family matrix factorisation.
Outcome: a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations