Challenge: Existing methods for cognate identification are based on distributions of phonemes and make little use of cognacy labels.
Approach: They propose a transformer-based architecture inspired by computational biology for automated cognate detection.
Outcome: The proposed architecture performs better than existing methods with increased supervision.

Similar Papers

Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction (2023.emnlp-main)

Copied to clipboard

Challenge: Phonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages.
Approach: They propose to use a protein language model to train on multiple sequence alignments to train a model on phonological reconstruction.
Outcome: The proposed model outperforms existing models on cognate reflex prediction task.
An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics (P19-1)

Copied to clipboard

Challenge: Existing methods for phylogenetic reconstruction of large datasets require time and computational power.
Approach: They propose a workflow for phylogenetic reconstruction on large datasets using two methods . they use a method for fast detection of cognates and a Bayesian method for inference . their results show that the methods take less than a few minutes to process language families .
Outcome: The proposed methods are fast and easy to use and close to gold standard cognate judgments and expert language family trees.
Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics? (N18-2)

Copied to clipboard

Challenge: Phylogenetic trees are hypotheses of how sets of related languages evolved in time.
Approach: They compare the performance of automatic cognate detection algorithms to classical manually annotated cognate sets.
Outcome: The proposed methods perform better than classically annotated cognate sets . future work on phylogenetic reconstruction can profit from the results .
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)

Copied to clipboard

Challenge: a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages .
Approach: They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation .
Outcome: The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU .
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability.
Approach: They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology.
Outcome: The proposed method achieves 94% accuracy on two pairs of Romance languages.
Cognition-aware Cognate Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to cognate detection use orthographic, phonetic and semantic similarity based features sets.
Approach: They propose a method for enriching feature sets with cognitive features extracted from gaze behaviour data from human readers’ gaze behaviour.
Outcome: The proposed method improves cognate detection performance by 10% and 12% over existing methods.
Probing Multilingual Cognate Prediction Models (2022.findings-acl)

Copied to clipboard

Challenge: linguistic interpretations of cognate prediction have been based on external analysis (accuracy, raw results, errors).
Approach: They propose to use character-based machine translation models to store linguistic and diachronic information but not in previously assumed ways.
Outcome: The proposed model stores linguistic and diachronic information but does not achieve it in previously assumed ways.
CogNet: A Large-Scale Cognate Database (P19-1)

Copied to clipboard

Challenge: Existing cognate databases have limited practical applications for research, despite their wide coverage and limited use in lexical tasks.
Approach: They introduce a new large-scale lexical database that provides cognates across languages.
Outcome: The proposed database contains 3.1 million cognate pairs across 338 languages and has an accuracy of 94%.
Identifying Cognates in English-Dutch and French-Dutch by means of Orthographic Information and Cross-lingual Word Embeddings (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to identify cognate pairs in English-Dutch and French-Dutsch combine orthographic information with cross-lingual word embeddings.
Approach: They combine traditional orthographic information with cross-lingual word embeddings to identify cognate pairs in English-Dutch and French-Dutsch.
Outcome: The proposed classifier achieves good results on the basis of orthographic information but improves by including semantic information in the form of cross-lingual word embeddings.
Can Cognate Prediction Be Modelled as a Low-Resource Machine Translation Task? (2021.findings-acl)

Copied to clipboard

Challenge: Existing work on cognate prediction based on similarities of two languages has not studied their differences or optimized architectural choices.
Approach: They compare statistical and neural MT architectures to a bilingual setup to test their hypothesis . they use monolingual pretraining, backtranslation and multilinguality to test the hypothesis based on the results .
Outcome: The proposed architectures can be used to generate cognates in a given language . the proposed architecture can be employed with monolingual pretraining, backtranslation and multilinguality .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations