Papers by Claudia Borg

10 papers
Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Maltese is a Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English.
Approach: They investigate whether Arabic-language resources can support Maltese natural language processing . they introduce transliteration schemes and machine translation approaches to align Arabic text with Maltesen .
Outcome: The proposed techniques can significantly improve Maltese natural language processing tasks.
COMET for Low-Resource Machine Translation Evaluation: A Case Study of English-Maltese and Spanish-Basque (2024.lrec-main)

Copied to clipboard

Challenge: Trainable metrics for machine translation evaluation have been scoring the highest correlations with human judgements in the meta-evaluations.
Approach: They run a crowd-based evaluation campaign to evaluate COMET-22 and fine-tune it to improve its performance.
Outcome: The proposed system outperforms BLEU and other lexical overlap metrics in the meta-evaluations.
Towards a Corpus of Spoken Maltese: Korpus tal-Malti Mitkellem, KMM (2024.lrec-main)

Copied to clipboard

Challenge: 'Corpus of Spoken Maltese' is a spoken corpus of spoken Malteser based on a gold-standard Core collection . initial results show that the ASR is robust enough to generate first-pass texts for annotators to work on, thus reducing the human effort and consequently, the cost involved.
Approach: They propose to create a “dedicated” spoken corpus of Maltese based on a gold-standard Core collection and a qualitative analysis of the output of a Malteser ASR system.
Outcome: The proposed corpus is based on the concept of a gold-standard Core collection and compares to human annotations.
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Maltese is the national language of Malta and is spoken by approximately 500,000 people.
Approach: They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment.
Outcome: The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta.
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)

Copied to clipboard

Challenge: Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones.
Approach: They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises .
Outcome: The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs.
Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-Switching (2024.eacl-long)

Copied to clipboard

Challenge: Multilingual models exhibit impressive cross-lingual transfer capabilities on unseen languages, but performance is impacted when there is a script disparity with the languages used in the model’s pre-training data.
Approach: They propose a novel method to align a resource-rich language's script with a target language and train a classifier that can make informed decisions regarding the appropriate processing of each token.
Outcome: The proposed model can be used to transfer a language's scripts across multiple languages, but it is suboptimal for mixed languages, where only a subset benefits while the rest is impeded.
Topic Classification and Headline Generation for Maltese Using a Public News Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for low-resource languages lack labelled data . public datasets only cover low-level syntactic tasks .
Approach: They propose to use a news tag multi-label classification and a summary task by generating its title to generate a new semantic dataset for Maltese.
Outcome: The proposed datasets show that current models lack the knowledge required to solve such tasks.
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)

Copied to clipboard

Challenge: a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images .
Approach: They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions .
Outcome: The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces.
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable performance across various NLP tasks, largely due to their generalisability and ability to perform tasks without additional training.
Approach: They evaluate the performance of 55 publicly available Large Language Models on Maltese, a low-resource language, using a newly introduced benchmark covering 11 discriminative and generative tasks.
Outcome: The proposed models perform poorly on discriminative and generative tasks and smaller fine-tuned models perform better across all tasks.
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America.
Approach: They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts.
Outcome: The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations