Ab Initio: Automatic Latin Proto-word Reconstruction (C18-1)

Copied to clipboard

Challenge: Existing methods for proto-word reconstruction are time-consuming and manual, but few studies have done it . a recent study used cognates to reconstruct ancient languages from their modern counterparts .
Approach: They propose to use Latin proto-words to automate the process of proto-language reconstruction . they leverage information from all modern languages and use conditional random fields for sequence labeling .
Outcome: The proposed method improves on previous results and requires less data . it is based on word forms in multiple Romance languages and on recurrent neural networks .

Similar Papers

Ab Antiquo: Neural Proto-language Reconstruction (2021.naacl-main)

Copied to clipboard

Challenge: Historical linguists have identified regularities in the process of historic sound change.
Approach: They propose a method to reconstruct proto-words based on cognates in daughter languages . they use a dataset of 8,000 comparative entries to analyze phonological changes .
Outcome: The proposed method outperforms conventional methods in a proto-word reconstruction task.
Neural Unsupervised Reconstruction of Protolanguage Word Forms (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for reconstructing ancient word forms use expectation-maximization . past work has used this method to predict simple phonological changes .
Approach: They extend expectation-maximization to predict phonological changes between ancient word forms and their cognates in modern languages.
Outcome: The proposed model reduces edit distance from the target word forms compared to previous methods.
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for protoword reconstruction are limited to a few languages.
Approach: They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it.
Outcome: The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features.
Transformed Protoform Reconstruction (2023.acl-short)

Copied to clipboard

Challenge: Historical linguists reconstruct proto-languages by identifying systematic sound changes that can be inferred from correspondences between attested daughter languages.
Approach: They propose to update their Latin protoform reconstruction model with the Transformer . romance data of 8,000 cognates spanning 5 languages and Chinese dataset are outperformed .
Outcome: The proposed model outperforms previous models on Romance and Chinese datasets.
Automatic Reconstruction of Missing Romanian Cognates and Unattested Latin Words (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for producing related words are based on sequence labeling .
Approach: They propose a method for producing related words based on sequence labeling . they aim to fill in gaps in incomplete cognate sets in Romance languages with Latin etymology and reconstruct uncertified Latin words.
Outcome: The proposed method fills in gaps in incomplete cognate sets in Romance languages with Latin etymology and reconstructs uncertified Latin words.
Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to discriminate between inherited and borrowed Latin words have been used to investigate the problem of automatic discrimination between a language's sound shifts.
Approach: They propose a new dataset to investigate the problem of automatically discriminating between inherited and borrowed Latin words in Romance languages.
Outcome: The proposed model can automatically discriminate between inherited and borrowed Latin words on two versions of the dataset, orthographic and phonetic.
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin (2024.lrec-main)

Copied to clipboard

Challenge: Historical linguists hypothesize systems of sound change to explain the evolution of language over time, but the evidence is limited.
Approach: They propose a dataset that consists of roughly 3,000 pairs of forms from Proto-Italic and Latin.
Outcome: The proposed dataset enables historical linguists to enhance other datasets by enhancing them with the existing datasets.
Improved Neural Protoform Reconstruction via Reflex Prediction (2024.lrec-main)

Copied to clipboard

Challenge: comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach .
Approach: They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model.
Outcome: The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets.
Simulating Language Evolution: a Tool for Historical Linguistics (C18-2)

Copied to clipboard

Challenge: Language change across space and time is one of the main concerns in historical linguistics.
Approach: They propose a web-based tool for word form production that predicts how a word evolves in a target language.
Outcome: The proposed method is language-agnostic and does not use external knowledge except for the training word pairs.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations