Papers by Søren Wichmann
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to detect cross-linguistic associations are not effective, but their effects are minor. |
| Approach: | They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon. |
| Outcome: | The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average). |
Towards identifying the optimal datasize for lexically-based Bayesian inference of linguistic phylogenies (C18-1)
Copied to clipboard
| Challenge: | Phylogenetic methods are used for linguistic phylogenies based on cognate matrices for words referring to a fix set of meanings. |
| Approach: | They propose to compute the quartet distance between the most stable meaning and the most unstable meaning . they rank meanings by stability and then compute the optimal number of meanings . |
| Outcome: | The proposed method is based on a set of language families with a fixed set of meanings. |