Papers by Johann-Mattis List
From Isolates to Families: Using Neural Networks for Automated Language Affiliation (2025.acl-long)
Copied to clipboard
| Challenge: | linguistic affiliation of languages to a common language family is traditionally carried out manually . large-scale standardized collections of multilingual wordlists and grammatical language structures could improve this . |
| Approach: | They propose to use lexical and grammatical data to classify languages into families using neural network models. |
| Outcome: | The proposed models outperform models trained on lexical and grammatical data while combining both types of data yields even better performance. |
Detecting Lexical Borrowings from Dominant Languages in Multilingual Wordlists (2023.eacl-main)
Copied to clipboard
| Challenge: | Language contact is reflected in the transfer of words from donor to recipient languages. |
| Approach: | They propose to use two classical sequence comparison methods and one machine learning method to detect lexical borrowings in contact situations where dominant languages play an important role. |
| Outcome: | The proposed methods outperform classical methods on a sample of seven Latin American languages. |
An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics (P19-1)
Copied to clipboard
| Challenge: | Existing methods for phylogenetic reconstruction of large datasets require time and computational power. |
| Approach: | They propose a workflow for phylogenetic reconstruction on large datasets using two methods . they use a method for fast detection of cognates and a Bayesian method for inference . their results show that the methods take less than a few minutes to process language families . |
| Outcome: | The proposed methods are fast and easy to use and close to gold standard cognate judgments and expert language family trees. |
Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics? (N18-2)
Copied to clipboard
| Challenge: | Phylogenetic trees are hypotheses of how sets of related languages evolved in time. |
| Approach: | They compare the performance of automatic cognate detection algorithms to classical manually annotated cognate sets. |
| Outcome: | The proposed methods perform better than classically annotated cognate sets . future work on phylogenetic reconstruction can profit from the results . |
Partial Colexifications Improve Concept Embeddings (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for embedding words from colexification networks are limited to the word level, ignoring lexical relations that would only hold for parts of words in a given language. |
| Approach: | They propose to embed concepts from automatically constructed colexification networks . they use lexical similarity ratings and word association data to evaluate the methods . |
| Outcome: | The proposed methods capture and represent different semantic relationships between concepts. |
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)
Copied to clipboard
| Challenge: | despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested. |
| Approach: | They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs. |
| Outcome: | The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code. |
Linguistic Survey of India and Polyglotta Africana: Two Retrostandardized Digital Editions of Large Historical Collections of Multilingual Wordlists (2024.lrec-main)
Copied to clipboard
| Challenge: | Linguistic Survey of India and Polyglotta Africana are two of the largest historical collections of multilingual wordlists. |
| Approach: | They present a retro-standardized edition of the Linguistic Survey of India and the Polyglotta Africana, which are two of the largest historical collections of multilingual wordlists. |
| Outcome: | The LSI and PA are the largest historical collections of multilingual wordlists . but no editions in which the original data is presented in standardized form have been produced so far . |
First Steps Towards the Integration of Resources on Historical Glossing Traditions in the History of Chinese: A Collection of Standardized Fǎnqiè Spellings from the Guǎngyùn (2024.lrec-main)
Copied to clipboard
| Challenge: | fnqiè spellings in the Gungyùn provide pronunciations for more than 20000 characters . historical glossing practices are a crucial part of the pronunciation of ancient Chinese characters - but no integrated resources have been developed on the topic . |
| Approach: | They propose to standardize digital versions of fnqiè spellings in the Gungyùn, one of the early rhyme books in the history of Chinese, providing pronunciations for more than 20000 characters. |
| Outcome: | The proposed resource can predict historical spellings with high precision and shed light on ancient glossing practices. |