Papers by Samia Touileb

10 papers
NoReC: The Norwegian Review Corpus (L18-1)

Copied to clipboard

Challenge: The Norwegian Review Corpus is a dataset of full-text reviews from major news sources.
Approach: This paper presents the Norwegian Review Corpus, created for document-level sentiment analysis.
Outcome: The corpus comprises more than 35,000 full-text reviews from a range of different domains.
JSEEGraph: Joint Structured Event Extraction as Graph Parsing (2023.starsem-1)

Copied to clipboard

Challenge: Existing approaches model event extraction using simplified datasets or sequence-labeling-based encodings.
Approach: They propose a graph-based event extraction framework that explicitly encodes entities and events in a single semantic graph.
Outcome: The proposed framework can handle nested event structures and solve different IE tasks jointly.
Measuring Normative and Descriptive Biases in Language Models Using Census Data (2023.eacl-main)

Copied to clipboard

Challenge: a new study examines how gender-based distributions of occupations are reflected in pre-trained language models.
Approach: They propose a method to measure to what degree pre-trained language models are aligned to normative and descriptive occupational distributions.
Outcome: The proposed method is language independent and can be extended to other dimensions of census data and demographic variables.
Exploring the Effects of Negation and Grammatical Tense on Bias Probes (2022.aacl-short)

Copied to clipboard

Challenge: Existing studies on negation in language models have shown that it does not affect correlations between gendered-pronouns and occupations.
Approach: They propose to add negation to bias probes to alter the grammatical tense of verbs in bias probe and to aggregate results across tenses to better represent existing correlations.
Outcome: The proposed method does not alter correlations between gendered-pronouns and occupations, but altering the grammatical tense of verbs does.
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains.
Approach: They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic.
Outcome: The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes.
Named Entity Recognition without Labelled Data: A Weak Supervision Approach (2020.acl-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) performance often degrades when applied to target domains that differ from the texts observed during training.
Approach: They propose a method to learn NER models in the absence of labelled data through weak supervision by using a broad spectrum of labelling functions to automatically annotate texts from the target domain.
Outcome: The proposed approach improves on two English datasets and shows that it improves by 7 percentage points on entity-level F1 scores compared to an out-of-domain neural NER model.
NERDz: A Preliminary Dataset of Named Entities for Algerian (2022.aacl-short)

Copied to clipboard

Challenge: NER is a fundamental task in information extraction and natural language processing.
Approach: They propose to build a manually annotated Algerian vernacular dataset using a recent extension to the Algerian NArabizi Treebank.
Outcome: The proposed dataset is the first of its kind for the Algerian vernacular dialect.
The interplay between language similarity and script on a novel multi-layer Algerian dialect corpus (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on cross-lingual transfer between languages with similar typology and languages of different scripts.
Approach: They propose to annotate Algerian user-generated comments with parallel annotations . they also investigate the effect of script vs. language similarity in cross-lingual transfer .
Outcome: The proposed model fine-tunes multi-lingual models on Algerian language and scripts . it shows that script vs. language similarity is important for part-of-speech tagging and sentiment analysis .
NorDiaChange: Diachronic Semantic Change Dataset for Norwegian (2022.lrec-1)

Copied to clipboard

Challenge: NorDiaChange is the first dataset of diachronic semantic change on the lexical level for Norwegian.
Approach: They describe a manual annotation process for a new dataset of diachronic semantic change for Norwegian.
Outcome: The proposed dataset covers the time periods related to pre- and post-war events, oil and gas discovery in Norway, and technological developments.
EDEN: A Dataset for Event Detection in Norwegian News (2024.lrec-main)

Copied to clipboard

Challenge: EDEN is the first dataset annotated with event information at the sentence level for the Norwegian language.
Approach: They propose to annotate Norwegian news text and transcribed speech using ACE event schema.
Outcome: The proposed dataset is the first annotated dataset for Norwegian, with a language-specific annotation process.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations