Multilingual Gradient Word-Order Typology from Universal Dependencies (2024.eacl-short)
Copied to clipboard
| Challenge: | Existing typological databases, including WALS and Grambank, suffer from inconsistencies due to categorical format. |
| Approach: | They propose a new seed dataset that uses continuous-valued data instead of categorical data to better reflect the variability of language. |
| Outcome: | The proposed dataset can be easily adapted to generate data for a broader set of features and languages. |
Similar Papers
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars. |
| Approach: | They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources. |
| Outcome: | The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios. |
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)
Copied to clipboard
| Challenge: | Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing . |
| Approach: | They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser. |
| Outcome: | The proposed model can be trained on 40 languages with the help of typological features. |
data2lang2vec: Data Driven Typological Features Completion (2025.coling-main)
Copied to clipboard
| Challenge: | Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database. |
| Approach: | They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features. |
| Outcome: | The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms. |
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning. |
| Approach: | They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization. |
| Outcome: | The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors. |
Uncovering Probabilistic Implications in Typological Knowledge Bases (P19-1)
Copied to clipboard
| Challenge: | linguistic typology is concerned with mapping out the relationships between languages with structural and functional properties. |
| Approach: | They propose a computational model which identifies known and new linguistic universals and uncovers them worthy of further linguistic investigation. |
| Outcome: | The proposed model outperforms baselines and knowledge base baselines. |
Massively Multilingual Token-Based Typology Using the Parallel Bible Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools. |
| Approach: | They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank. |
| Outcome: | The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language. |
A Probabilistic Generative Model of Linguistic Typology (N19-1)
Copied to clipboard
| Challenge: | a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages. |
| Approach: | They propose a generative model of language based on exponential-family matrix factorisation. |
| Outcome: | a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features. |
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)
Copied to clipboard
| Challenge: | linguistic typology is the classification of languages according to their linguistic properties. |
| Approach: | They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale. |
| Outcome: | The proposed model can predict typological properties on a massively multilingual scale. |
What is ”Typological Diversity” in NLP? (2024.emnlp-main)
Copied to clipboard
| Challenge: | linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims. |
| Approach: | They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages. |
| Outcome: | The proposed measures show that skewed language selection can lead to overestimated multilingual performance. |