Papers by Meriem Beloucif

11 papers
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)

Copied to clipboard

Challenge: Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages.
Approach: They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets.
Outcome: The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers.
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages (2025.naacl-long)

Copied to clipboard

Challenge: Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated.
Approach: They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance.
Outcome: The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.
Which is Better for Deep Learning: Python or MATLAB? Answering Comparative Questions in Natural Language (2021.eacl-demos)

Copied to clipboard

Challenge: Comparative QA is a challenging task since it requires collecting evidence from many different sources.
Approach: They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices.
Outcome: The proposed system can be used in personal assistants, chatbots, and similar NLP devices.
Elvis vs. M. Jackson: Who has More Albums? Classification and Identification of Elements in Comparative Questions (2022.lrec-1)

Copied to clipboard

Challenge: Comparative Question Answering (cQA) is the task of providing accurate answers to questions . most question answering systems focus on answering factoid questions, but they fail at answering comparative questions in an efficient argumentative manner.
Approach: They propose two new open-domain datasets for identifying and labeling comparative questions . they use a binary classification task and an unsupervised sequence labeling task .
Outcome: The proposed datasets reach close-to-human results on a binary classification task with a neural model using ALBERT embeddings.
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
Probing Pre-trained Language Models for Semantic Attributes and their Values (2021.findings-emnlp)

Copied to clipboard

Challenge: Pretrained language models (PTLMs) are used for many tasks including syntax, semantics and commonsense.
Approach: They propose to integrate semantic attributes and their values into pretrained language models to improve their performance on many natural language processing tasks.
Outcome: The proposed model performs better on masked tokens than humans on this task.
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Africa has the highest linguistic diversity among all continents.
Approach: They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges .
Outcome: The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task .
Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: ZOPrO is a novel algorithm designed for *Preference Optimisation* in large language models.
Approach: They propose a ZO algorithm designed for *Preference Optimisation* in LLMs that uses function evaluations instead of gradients to reduce memory usage.
Outcome: The proposed method improves reward signals while achieving convergence times comparable to first-order methods.
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers.
Approach: They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch.
Outcome: The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages.
BERTie Bott’s Every Flavor Labels: A Tasty Introduction to Semantic Role Labeling for Galician (2023.emnlp-main)

Copied to clipboard

Challenge: Existing corpora, WordNet, and dependency parsing are used to build a semantic role labeling system.
Approach: They use existing corpora, WordNet, and dependency parsing to build a Galician dataset for training semantic role labeling systems.
Outcome: The proposed model outperforms the 2009 CoNLL Shared Task by 0.83 on Spanish datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations