| Challenge: | Existing methods for proto-word reconstruction are time-consuming and manual, but few studies have done it . a recent study used cognates to reconstruct ancient languages from their modern counterparts . |
| Approach: | They propose to use Latin proto-words to automate the process of proto-language reconstruction . they leverage information from all modern languages and use conditional random fields for sequence labeling . |
| Outcome: | The proposed method improves on previous results and requires less data . it is based on word forms in multiple Romance languages and on recurrent neural networks . |
Similar Papers
Ab Antiquo: Neural Proto-language Reconstruction (2021.naacl-main)
Copied to clipboard
| Challenge: | Historical linguists have identified regularities in the process of historic sound change. |
| Approach: | They propose a method to reconstruct proto-words based on cognates in daughter languages . they use a dataset of 8,000 comparative entries to analyze phonological changes . |
| Outcome: | The proposed method outperforms conventional methods in a proto-word reconstruction task. |
Neural Unsupervised Reconstruction of Protolanguage Word Forms (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for reconstructing ancient word forms use expectation-maximization . past work has used this method to predict simple phonological changes . |
| Approach: | They extend expectation-maximization to predict phonological changes between ancient word forms and their cognates in modern languages. |
| Outcome: | The proposed model reduces edit distance from the target word forms compared to previous methods. |
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for protoword reconstruction are limited to a few languages. |
| Approach: | They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it. |
| Outcome: | The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features. |
Transformed Protoform Reconstruction (2023.acl-short)
Copied to clipboard
| Challenge: | Historical linguists reconstruct proto-languages by identifying systematic sound changes that can be inferred from correspondences between attested daughter languages. |
| Approach: | They propose to update their Latin protoform reconstruction model with the Transformer . romance data of 8,000 cognates spanning 5 languages and Chinese dataset are outperformed . |
| Outcome: | The proposed model outperforms previous models on Romance and Chinese datasets. |
Automatic Reconstruction of Missing Romanian Cognates and Unattested Latin Words (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for producing related words are based on sequence labeling . |
| Approach: | They propose a method for producing related words based on sequence labeling . they aim to fill in gaps in incomplete cognate sets in Romance languages with Latin etymology and reconstruct uncertified Latin words. |
| Outcome: | The proposed method fills in gaps in incomplete cognate sets in Romance languages with Latin etymology and reconstructs uncertified Latin words. |
Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to discriminate between inherited and borrowed Latin words have been used to investigate the problem of automatic discrimination between a language's sound shifts. |
| Approach: | They propose a new dataset to investigate the problem of automatically discriminating between inherited and borrowed Latin words in Romance languages. |
| Outcome: | The proposed model can automatically discriminate between inherited and borrowed Latin words on two versions of the dataset, orthographic and phonetic. |
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin (2024.lrec-main)
Copied to clipboard
| Challenge: | Historical linguists hypothesize systems of sound change to explain the evolution of language over time, but the evidence is limited. |
| Approach: | They propose a dataset that consists of roughly 3,000 pairs of forms from Proto-Italic and Latin. |
| Outcome: | The proposed dataset enables historical linguists to enhance other datasets by enhancing them with the existing datasets. |
Improved Neural Protoform Reconstruction via Reflex Prediction (2024.lrec-main)
Copied to clipboard
| Challenge: | comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach . |
| Approach: | They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model. |
| Outcome: | The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets. |
Simulating Language Evolution: a Tool for Historical Linguistics (C18-2)
Copied to clipboard
| Challenge: | Language change across space and time is one of the main concerns in historical linguistics. |
| Approach: | They propose a web-based tool for word form production that predicts how a word evolves in a target language. |
| Outcome: | The proposed method is language-agnostic and does not use external knowledge except for the training word pairs. |
Advances in Pre-Training Distributed Word Representations (L18-1)
Copied to clipboard
| Challenge: | Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications. |
| Approach: | They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations. |
| Outcome: | The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data. |