Papers by Søren Wichmann

3 papers
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to detect cross-linguistic associations are not effective, but their effects are minor.
Approach: They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon.
Outcome: The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average).
Towards identifying the optimal datasize for lexically-based Bayesian inference of linguistic phylogenies (C18-1)

Copied to clipboard

Challenge: Phylogenetic methods are used for linguistic phylogenies based on cognate matrices for words referring to a fix set of meanings.
Approach: They propose to compute the quartet distance between the most stable meaning and the most unstable meaning . they rank meanings by stability and then compute the optimal number of meanings .
Outcome: The proposed method is based on a set of language families with a fixed set of meanings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations