Papers by Claudia Borg
Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Maltese is a Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. |
| Approach: | They investigate whether Arabic-language resources can support Maltese natural language processing . they introduce transliteration schemes and machine translation approaches to align Arabic text with Maltesen . |
| Outcome: | The proposed techniques can significantly improve Maltese natural language processing tasks. |
COMET for Low-Resource Machine Translation Evaluation: A Case Study of English-Maltese and Spanish-Basque (2024.lrec-main)
Copied to clipboard
| Challenge: | Trainable metrics for machine translation evaluation have been scoring the highest correlations with human judgements in the meta-evaluations. |
| Approach: | They run a crowd-based evaluation campaign to evaluate COMET-22 and fine-tune it to improve its performance. |
| Outcome: | The proposed system outperforms BLEU and other lexical overlap metrics in the meta-evaluations. |
Towards a Corpus of Spoken Maltese: Korpus tal-Malti Mitkellem, KMM (2024.lrec-main)
Copied to clipboard
| Challenge: | 'Corpus of Spoken Maltese' is a spoken corpus of spoken Malteser based on a gold-standard Core collection . initial results show that the ASR is robust enough to generate first-pass texts for annotators to work on, thus reducing the human effort and consequently, the cost involved. |
| Approach: | They propose to create a “dedicated” spoken corpus of Maltese based on a gold-standard Core collection and a qualitative analysis of the output of a Malteser ASR system. |
| Outcome: | The proposed corpus is based on the concept of a gold-standard Core collection and compares to human annotations. |
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)
Copied to clipboard
Carlos Daniel Hernandez Mena, Albert Gatt, Andrea DeMarco, Claudia Borg, Lonneke van der Plas, Amanda Muscat, Ian Padovani
| Challenge: | Maltese is the national language of Malta and is spoken by approximately 500,000 people. |
| Approach: | They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment. |
| Outcome: | The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta. |
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)
Copied to clipboard
Lionel Nicolas, Verena Lyding, Claudia Borg, Corina Forascu, Karën Fort, Katerina Zdravkova, Iztok Kosem, Jaka Čibej, Špela Arhar Holdt, Alice Millour, Alexander König, Christos Rodosthenous, Federico Sangati, Umair ul Hassan, Anisia Katinskaia, Anabela Barreiro, Lavinia Aparaschivei, Yaakov HaCohen-Kerner
| Challenge: | Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones. |
| Approach: | They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises . |
| Outcome: | The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs. |
Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-Switching (2024.eacl-long)
Copied to clipboard
| Challenge: | Multilingual models exhibit impressive cross-lingual transfer capabilities on unseen languages, but performance is impacted when there is a script disparity with the languages used in the model’s pre-training data. |
| Approach: | They propose a novel method to align a resource-rich language's script with a target language and train a classifier that can make informed decisions regarding the appropriate processing of each token. |
| Outcome: | The proposed model can be used to transfer a language's scripts across multiple languages, but it is suboptimal for mixed languages, where only a subset benefits while the rest is impeded. |
Topic Classification and Headline Generation for Maltese Using a Public News Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for low-resource languages lack labelled data . public datasets only cover low-level syntactic tasks . |
| Approach: | They propose to use a news tag multi-label classification and a summary task by generating its title to generate a new semantic dataset for Maltese. |
| Outcome: | The proposed datasets show that current models lack the knowledge required to solve such tasks. |
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)
Copied to clipboard
Albert Gatt, Marc Tanti, Adrian Muscat, Patrizia Paggio, Reuben A Farrugia, Claudia Borg, Kenneth P Camilleri, Michael Rosner, Lonneke van der Plas
| Challenge: | a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images . |
| Approach: | They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions . |
| Outcome: | The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces. |
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable performance across various NLP tasks, largely due to their generalisability and ability to perform tasks without additional training. |
| Approach: | They evaluate the performance of 55 publicly available Large Language Models on Maltese, a low-resource language, using a newly introduced benchmark covering 11 discriminative and generative tasks. |
| Outcome: | The proposed models perform poorly on discriminative and generative tasks and smaller fine-tuned models perform better across all tasks. |
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)
Copied to clipboard
Karen Fort, Laura Alonso Alemany, Luciana Benotti, Julien Bezançon, Claudia Borg, Marthese Borg, Yongjian Chen, Fanny Ducel, Yoann Dupont, Guido Ivetta, Zhijian Li, Margot Mieskes, Marco Naguib, Yuyan Qian, Matteo Radaelli, Wolfgang S. Schmeisser-Nieto, Emma Raimundo Schulz, Thiziri Saci, Sarah Saidi, Javier Torroba Marchante, Shilin Xie, Sergio E. Zanotto, Aurélie Névéol
| Challenge: | Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America. |
| Approach: | They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts. |
| Outcome: | The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories . |