Papers by Johann-Mattis List

8 papers
From Isolates to Families: Using Neural Networks for Automated Language Affiliation (2025.acl-long)

Copied to clipboard

Challenge: linguistic affiliation of languages to a common language family is traditionally carried out manually . large-scale standardized collections of multilingual wordlists and grammatical language structures could improve this .
Approach: They propose to use lexical and grammatical data to classify languages into families using neural network models.
Outcome: The proposed models outperform models trained on lexical and grammatical data while combining both types of data yields even better performance.
Detecting Lexical Borrowings from Dominant Languages in Multilingual Wordlists (2023.eacl-main)

Copied to clipboard

Challenge: Language contact is reflected in the transfer of words from donor to recipient languages.
Approach: They propose to use two classical sequence comparison methods and one machine learning method to detect lexical borrowings in contact situations where dominant languages play an important role.
Outcome: The proposed methods outperform classical methods on a sample of seven Latin American languages.
An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics (P19-1)

Copied to clipboard

Challenge: Existing methods for phylogenetic reconstruction of large datasets require time and computational power.
Approach: They propose a workflow for phylogenetic reconstruction on large datasets using two methods . they use a method for fast detection of cognates and a Bayesian method for inference . their results show that the methods take less than a few minutes to process language families .
Outcome: The proposed methods are fast and easy to use and close to gold standard cognate judgments and expert language family trees.
Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics? (N18-2)

Copied to clipboard

Challenge: Phylogenetic trees are hypotheses of how sets of related languages evolved in time.
Approach: They compare the performance of automatic cognate detection algorithms to classical manually annotated cognate sets.
Outcome: The proposed methods perform better than classically annotated cognate sets . future work on phylogenetic reconstruction can profit from the results .
Partial Colexifications Improve Concept Embeddings (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding words from colexification networks are limited to the word level, ignoring lexical relations that would only hold for parts of words in a given language.
Approach: They propose to embed concepts from automatically constructed colexification networks . they use lexical similarity ratings and word association data to evaluate the methods .
Outcome: The proposed methods capture and represent different semantic relationships between concepts.
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)

Copied to clipboard

Challenge: despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested.
Approach: They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs.
Outcome: The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code.
Linguistic Survey of India and Polyglotta Africana: Two Retrostandardized Digital Editions of Large Historical Collections of Multilingual Wordlists (2024.lrec-main)

Copied to clipboard

Challenge: Linguistic Survey of India and Polyglotta Africana are two of the largest historical collections of multilingual wordlists.
Approach: They present a retro-standardized edition of the Linguistic Survey of India and the Polyglotta Africana, which are two of the largest historical collections of multilingual wordlists.
Outcome: The LSI and PA are the largest historical collections of multilingual wordlists . but no editions in which the original data is presented in standardized form have been produced so far .
First Steps Towards the Integration of Resources on Historical Glossing Traditions in the History of Chinese: A Collection of Standardized Fǎnqiè Spellings from the Guǎngyùn (2024.lrec-main)

Copied to clipboard

Challenge: fnqiè spellings in the Gungyùn provide pronunciations for more than 20000 characters . historical glossing practices are a crucial part of the pronunciation of ancient Chinese characters - but no integrated resources have been developed on the topic .
Approach: They propose to standardize digital versions of fnqiè spellings in the Gungyùn, one of the early rhyme books in the history of Chinese, providing pronunciations for more than 20000 characters.
Outcome: The proposed resource can predict historical spellings with high precision and shed light on ancient glossing practices.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations