Papers by Ioan-Bogdan Iordache
Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | lexical divergence between cognate and borrowings is studied in the five Romance languages. |
| Approach: | They propose to use etymological dictionaries to extract deceptive cognates and borrowings automatically based on usage and freely publish the lexicon of obtained true and deceptives in every Romance language pair. |
| Outcome: | The proposed algorithms are based on the most complete and reliable dataset of cognate words based etymological dictionaries for the five main Romance languages. |
Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages (2024.lrec-main)
Copied to clipboard
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for discriminating between cognates and borrowings are difficult, but they provide a deeper insight into the history of a language and allow for a better characterization of language relatedness. |
| Approach: | They propose a computational approach for discriminating between cognates and borrowings based on a comprehensive database of Romance cognates. |
| Outcome: | The proposed approach is the most comprehensive in terms of covered languages. |
Detecting Optimism in Tweets using Knowledge Distillation and Linguistic Analysis of Optimism (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study has established sentiment analysis as an alluring problem, but many feelings are left unexplored. |
| Approach: | They propose a framework to learn the polarity of emotions from Twitter posts . they compare optimism detection with sentiment analysis and hate speech detection . |
| Outcome: | The proposed framework differs between optimistic and pessimistic users on the Optimism/Pessimism Twitter dataset. |
RoCode: A Dataset for Measuring Code Intelligence from Problem Definitions in Romanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models are capable of solving tasks in natural language, but most tests assume they are written in English. |
| Approach: | They propose to use a dataset to measure the generalization power of large language models in a language other than English to evaluate their code intelligence. |
| Outcome: | The proposed dataset provides a benchmark for evaluating the code intelligence of language models trained on Romanian / multilingual text and a fine-tuning set for pretrained Romanian models. |
A Computational Exploration of Pejorative Language in Social Media (2021.findings-emnlp)
Copied to clipboard
| Challenge: | In this paper, we examine the problem of pejorative language, an under-explored topic in computational linguistics. |
| Approach: | They propose to automatically disambiguate pejorative usage in social media . they leverage online dictionaries to build a multilingual lexicon of pejorativ terms . |
| Outcome: | The proposed model can automatically disambiguate pejorative usage in social media posts . the proposed model is based on dictionaries and tweets . |
Verba volant, scripta volant? Don’t worry! There are computational solutions for protoword reconstruction (2024.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing methods for protoword reconstruction are limited to a few languages. |
| Approach: | They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it. |
| Outcome: | The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features. |
Investigating the Relationship Between Romanian Financial News and Closing Prices from the Bucharest Stock Exchange (2022.lrec-1)
Copied to clipboard
| Challenge: | a new data set is used to extract information related to one company . a model that is based on previous information about transactions is not enough . |
| Approach: | They use a Romanian financial news website to extract only information related to one company . they use lexicon-based Vader tool, Financial BERT and Transformer-based models . |
| Outcome: | The proposed model shows that the extracted sentiment scores correlate with stock closing prices . the proposed model is based on data from a Romanian financial news website . |
It takes two to borrow: a donor and a recipient. Who’s who? (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for identifying the direction of borrowing are limited. |
| Approach: | They propose strong benchmarks for automatic borrowing direction detection by using a borrowings dataset from the recent RoBoCoP database for five Romance languages. |
| Outcome: | The proposed model improves the accuracy of the proposed task and proposes additional directions for future work. |
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification (2023.emnlp-main)
Copied to clipboard
Liviu Dinu, Ana Uban, Alina Cristea, Anca Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
| Challenge: | Existing databases for romance cognates are scattered, incomplete, noisy, or have uncertain availability. |
| Approach: | They propose to use etymological information to identify Romance cognates and borrowings from dictionaries to identify their ethymology. |
| Outcome: | The proposed method achieves 94% accuracy on two pairs of Romance languages. |