Papers with WALS

5 papers
Multilingual Gradient Word-Order Typology from Universal Dependencies (2024.eacl-short)

Copied to clipboard

Challenge: Existing typological databases, including WALS and Grambank, suffer from inconsistencies due to categorical format.
Approach: They propose a new seed dataset that uses continuous-valued data instead of categorical data to better reflect the variability of language.
Outcome: The proposed dataset can be easily adapted to generate data for a broader set of features and languages.
Does Typological Blinding Impede Cross-Lingual Sharing? (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on bridging the performance gap between high- and low-resource languages has only found minor benefits from using typological information.
Approach: They propose to use typological features to train models in a cross-lingual setting to learn latent weights between languages.
Outcome: The proposed model overshadows the utility of explicitly using typological features by ignoring them, and shows that encouraging sharing according to typology improves performance.
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars.
Approach: They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources.
Outcome: The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
Language Embeddings for Typology and Cross-lingual Transfer Learning (2021.acl-long)

Copied to clipboard

Challenge: Recent efforts to leverage multilingual datasets highlight potential of multilingual models that can perform well across various languages.
Approach: They propose to generate language representations that capture relationships among languages and evaluate them using WALS and two extrinsic tasks.
Outcome: The proposed model can be leveraged in cross-lingual tasks without parallel data . the proposed model is based on the World Atlas of Language Structures (WALS) and two extrinsic tasks .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations