Challenge: Phylogenetic trees are hypotheses of how sets of related languages evolved in time.
Approach: They compare the performance of automatic cognate detection algorithms to classical manually annotated cognate sets.
Outcome: The proposed methods perform better than classically annotated cognate sets . future work on phylogenetic reconstruction can profit from the results .

Similar Papers

An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics (P19-1)

Copied to clipboard

Challenge: Existing methods for phylogenetic reconstruction of large datasets require time and computational power.
Approach: They propose a workflow for phylogenetic reconstruction on large datasets using two methods . they use a method for fast detection of cognates and a Bayesian method for inference . their results show that the methods take less than a few minutes to process language families .
Outcome: The proposed methods are fast and easy to use and close to gold standard cognate judgments and expert language family trees.
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for protoword reconstruction are limited to a few languages.
Approach: They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it.
Outcome: The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features.
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness.
Approach: They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates.
Outcome: The proposed approach is the most comprehensive in terms of covered languages.
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)

Copied to clipboard

Challenge: a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages .
Approach: They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation .
Outcome: The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU .
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability.
Approach: They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology.
Outcome: The proposed method achieves 94% accuracy on two pairs of Romance languages.
CogNet: A Large-Scale Cognate Database (P19-1)

Copied to clipboard

Challenge: Existing cognate databases have limited practical applications for research, despite their wide coverage and limited use in lexical tasks.
Approach: They introduce a new large-scale lexical database that provides cognates across languages.
Outcome: The proposed database contains 3.1 million cognate pairs across 338 languages and has an accuracy of 94%.
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods for cognate identification are based on distributions of phonemes and make little use of cognacy labels.
Approach: They propose a transformer-based architecture inspired by computational biology for automated cognate detection.
Outcome: The proposed architecture performs better than existing methods with increased supervision.
Probing Multilingual Cognate Prediction Models (2022.findings-acl)

Copied to clipboard

Challenge: linguistic interpretations of cognate prediction have been based on external analysis (accuracy, raw results, errors).
Approach: They propose to use character-based machine translation models to store linguistic and diachronic information but not in previously assumed ways.
Outcome: The proposed model stores linguistic and diachronic information but does not achieve it in previously assumed ways.
Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of Ge’ez, Tigrinya, and Amharic (2025.coling-main)

Copied to clipboard

Challenge: As languages evolve, we risk losing ancestral languages.
Approach: They propose to use cognates to reconstruct proto-languages from cognates in child languages that have likely evolved from the same word in the proto-linguistics.
Outcome: The proposed method is based on automatic cognate detection and in-context learning with GPT-4o to generate the proto-language from the cognates and use Sequence-to-Sequence models.
Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction (2023.emnlp-main)

Copied to clipboard

Challenge: Phonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages.
Approach: They propose to use a protein language model to train on multiple sequence alignments to train a model on phonological reconstruction.
Outcome: The proposed model outperforms existing models on cognate reflex prediction task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations