Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics? (N18-2)
Copied to clipboard
| Challenge: | Phylogenetic trees are hypotheses of how sets of related languages evolved in time. |
| Approach: | They compare the performance of automatic cognate detection algorithms to classical manually annotated cognate sets. |
| Outcome: | The proposed methods perform better than classically annotated cognate sets . future work on phylogenetic reconstruction can profit from the results . |
Similar Papers
An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics (P19-1)
Copied to clipboard
| Challenge: | Existing methods for phylogenetic reconstruction of large datasets require time and computational power. |
| Approach: | They propose a workflow for phylogenetic reconstruction on large datasets using two methods . they use a method for fast detection of cognates and a Bayesian method for inference . their results show that the methods take less than a few minutes to process language families . |
| Outcome: | The proposed methods are fast and easy to use and close to gold standard cognate judgments and expert language family trees. |
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for protoword reconstruction are limited to a few languages. |
| Approach: | They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it. |
| Outcome: | The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features. |
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)
Copied to clipboard
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness. |
| Approach: | They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates. |
| Outcome: | The proposed approach is the most comprehensive in terms of covered languages. |
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)
Copied to clipboard
Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Anca Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability. |
| Approach: | They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology. |
| Outcome: | The proposed method achieves 94% accuracy on two pairs of Romance languages. |
CogNet: A Large-Scale Cognate Database (P19-1)
Copied to clipboard
| Challenge: | Existing cognate databases have limited practical applications for research, despite their wide coverage and limited use in lexical tasks. |
| Approach: | They introduce a new large-scale lexical database that provides cognates across languages. |
| Outcome: | The proposed database contains 3.1 million cognate pairs across 338 languages and has an accuracy of 94%. |
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing methods for cognate identification are based on distributions of phonemes and make little use of cognacy labels. |
| Approach: | They propose a transformer-based architecture inspired by computational biology for automated cognate detection. |
| Outcome: | The proposed architecture performs better than existing methods with increased supervision. |
Probing Multilingual Cognate Prediction Models (2022.findings-acl)
Copied to clipboard
| Challenge: | linguistic interpretations of cognate prediction have been based on external analysis (accuracy, raw results, errors). |
| Approach: | They propose to use character-based machine translation models to store linguistic and diachronic information but not in previously assumed ways. |
| Outcome: | The proposed model stores linguistic and diachronic information but does not achieve it in previously assumed ways. |
Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of Ge’ez, Tigrinya, and Amharic (2025.coling-main)
Copied to clipboard
| Challenge: | As languages evolve, we risk losing ancestral languages. |
| Approach: | They propose to use cognates to reconstruct proto-languages from cognates in child languages that have likely evolved from the same word in the proto-linguistics. |
| Outcome: | The proposed method is based on automatic cognate detection and in-context learning with GPT-4o to generate the proto-language from the cognates and use Sequence-to-Sequence models. |
Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction (2023.emnlp-main)
Copied to clipboard
| Challenge: | Phonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages. |
| Approach: | They propose to use a protein language model to train on multiple sequence alignments to train a model on phonological reconstruction. |
| Outcome: | The proposed model outperforms existing models on cognate reflex prediction task. |