Papers by Lucia Pitarch
Building MUSCLE, a Dataset for MUltilingual Semantic Classification of Links between Entities (2024.lrec-main)
Copied to clipboard
| Challenge: | In this paper we present a dataset for MUltilingual Lexical Relation Classification (LRC) systems with 27K pairs of universal concepts selected from Wikidata, a large and highly multilingual factual Knowledge Graph (KG). |
| Approach: | They propose a dataset for MUltilingual lexico-semantic Classification of Links between Entities using 27K pairs of universal concepts selected from Wikidata. |
| Outcome: | The proposed dataset bridges lexical and conceptual semantics, avoids linguistic memorization, is domain-balanced across entities, and enables enrichment and hierarchical information retrieval. |
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)
Copied to clipboard
Dagmar Gromann, Hugo Goncalo Oliveira, Lucia Pitarch, Elena-Simona Apostol, Jordi Bernad, Eliot Bytyçi, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabik, Jorge Gracia, Letizia Granata, Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia Di Buono, Ana Ostroški Anić, Sigita Rackevičienė, Ricardo Rodrigues, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stanković, Ciprian-Octavian Truică, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
| Challenge: | Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions. |
| Approach: | They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge. |
| Outcome: | The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian. |
No clues good clues: out of context Lexical Relation Classification (2023.acl-long)
Copied to clipboard
| Challenge: | Pre-trained language models (PTLMs) are used to predict lexical relations between words. |
| Approach: | They propose to use pre-trained language models to fine-tune and exploit verbalized text for linguistically motivated tasks. |
| Outcome: | The proposed model outperforms graded Lexical Entailment and lexical relation classification with very simple prompts. |