Challenge: Existing typological databases, including WALS and Grambank, suffer from inconsistencies due to categorical format.
Approach: They propose a new seed dataset that uses continuous-valued data instead of categorical data to better reflect the variability of language.
Outcome: The proposed dataset can be easily adapted to generate data for a broader set of features and languages.

Similar Papers

Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars.
Approach: They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources.
Outcome: The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios.
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)

Copied to clipboard

Challenge: Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing .
Approach: They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser.
Outcome: The proposed model can be trained on 40 languages with the help of typological features.
data2lang2vec: Data Driven Typological Features Completion (2025.coling-main)

Copied to clipboard

Challenge: Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database.
Approach: They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features.
Outcome: The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms.
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning.
Approach: They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization.
Outcome: The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors.
Uncovering Probabilistic Implications in Typological Knowledge Bases (P19-1)

Copied to clipboard

Challenge: linguistic typology is concerned with mapping out the relationships between languages with structural and functional properties.
Approach: They propose a computational model which identifies known and new linguistic universals and uncovers them worthy of further linguistic investigation.
Outcome: The proposed model outperforms baselines and knowledge base baselines.
Massively Multilingual Token-Based Typology Using the Parallel Bible Corpus (2024.lrec-main)

Copied to clipboard

Challenge: linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools.
Approach: They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank.
Outcome: The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language.
A Probabilistic Generative Model of Linguistic Typology (N19-1)

Copied to clipboard

Challenge: a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages.
Approach: They propose a generative model of language based on exponential-family matrix factorisation.
Outcome: a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
What is ”Typological Diversity” in NLP? (2024.emnlp-main)

Copied to clipboard

Challenge: linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims.
Approach: They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages.
Outcome: The proposed measures show that skewed language selection can lead to overestimated multilingual performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations