Papers by Simona Georgescu

6 papers
Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to discriminate between inherited and borrowed Latin words have been used to investigate the problem of automatic discrimination between a language's sound shifts.
Approach: They propose a new dataset to investigate the problem of automatically discriminating between inherited and borrowed Latin words in Romance languages.
Outcome: The proposed model can automatically discriminate between inherited and borrowed Latin words on two versions of the dataset, orthographic and phonetic.
Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages (2025.emnlp-main)

Copied to clipboard

Challenge: lexical divergence between cognate and borrowings is studied in the five Romance languages.
Approach: They propose to use etymological dictionaries to extract deceptive cognates and borrowings automatically based on usage and freely publish the lexicon of obtained true and deceptives in every Romance language pair.
Outcome: The proposed algorithms are based on the most complete and reliable dataset of cognate words based etymological dictionaries for the five main Romance languages.
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness.
Approach: They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates.
Outcome: The proposed approach is the most comprehensive in terms of covered languages.
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for protoword reconstruction are limited to a few languages.
Approach: They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it.
Outcome: The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features.
It takes two to borrow: a donor and a recipient. Who’s who? (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for identifying the direction of borrowing are limited.
Approach: They propose strong benchmarks for automatic borrowing direction detection by using a borrowings dataset from the recent RoBoCoP database for five Romance languages.
Outcome: The proposed model improves the accuracy of the proposed task and proposes additional directions for future work.
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability.
Approach: They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology.
Outcome: The proposed method achieves 94% accuracy on two pairs of Romance languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations