Papers by Simona Georgescu
Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to discriminate between inherited and borrowed Latin words have been used to investigate the problem of automatic discrimination between a language's sound shifts. |
| Approach: | They propose a new dataset to investigate the problem of automatically discriminating between inherited and borrowed Latin words in Romance languages. |
| Outcome: | The proposed model can automatically discriminate between inherited and borrowed Latin words on two versions of the dataset, orthographic and phonetic. |
Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | lexical divergence between cognate and borrowings is studied in the five Romance languages. |
| Approach: | They propose to use etymological dictionaries to extract deceptive cognates and borrowings automatically based on usage and freely publish the lexicon of obtained true and deceptives in every Romance language pair. |
| Outcome: | The proposed algorithms are based on the most complete and reliable dataset of cognate words based etymological dictionaries for the five main Romance languages. |
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)
Copied to clipboard
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness. |
| Approach: | They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates. |
| Outcome: | The proposed approach is the most comprehensive in terms of covered languages. |
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for protoword reconstruction are limited to a few languages. |
| Approach: | They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it. |
| Outcome: | The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features. |
It takes two to borrow: a donor and a recipient. Who’s who? (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for identifying the direction of borrowing are limited. |
| Approach: | They propose strong benchmarks for automatic borrowing direction detection by using a borrowings dataset from the recent RoBoCoP database for five Romance languages. |
| Outcome: | The proposed model improves the accuracy of the proposed task and proposes additional directions for future work. |
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Anca Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability. |
| Approach: | They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology. |
| Outcome: | The proposed method achieves 94% accuracy on two pairs of Romance languages. |