Papers by Samia Touileb
NoReC: The Norwegian Review Corpus (L18-1)
Copied to clipboard
Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, Fredrik Jørgensen
| Challenge: | The Norwegian Review Corpus is a dataset of full-text reviews from major news sources. |
| Approach: | This paper presents the Norwegian Review Corpus, created for document-level sentiment analysis. |
| Outcome: | The corpus comprises more than 35,000 full-text reviews from a range of different domains. |
JSEEGraph: Joint Structured Event Extraction as Graph Parsing (2023.starsem-1)
Copied to clipboard
| Challenge: | Existing approaches model event extraction using simplified datasets or sequence-labeling-based encodings. |
| Approach: | They propose a graph-based event extraction framework that explicitly encodes entities and events in a single semantic graph. |
| Outcome: | The proposed framework can handle nested event structures and solve different IE tasks jointly. |
Measuring Normative and Descriptive Biases in Language Models Using Census Data (2023.eacl-main)
Copied to clipboard
| Challenge: | a new study examines how gender-based distributions of occupations are reflected in pre-trained language models. |
| Approach: | They propose a method to measure to what degree pre-trained language models are aligned to normative and descriptive occupational distributions. |
| Outcome: | The proposed method is language independent and can be extended to other dimensions of census data and demographic variables. |
Exploring the Effects of Negation and Grammatical Tense on Bias Probes (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing studies on negation in language models have shown that it does not affect correlations between gendered-pronouns and occupations. |
| Approach: | They propose to add negation to bias probes to alter the grammatical tense of verbs in bias probe and to aggregate results across tenses to better represent existing correlations. |
| Outcome: | The proposed method does not alter correlations between gendered-pronouns and occupations, but altering the grammatical tense of verbs does. |
Identifying Sentiments in Algerian Code-switched User-generated Comments (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on sentiment analysis for the Arabic variety, but it has been extended to other domains. |
| Approach: | They build a corpus of 36,000 code-switched user-generated comments annotated for sentiments in Algerian Arabic. |
| Outcome: | The proposed model performs better on unedited code-switched and unbalanced data across sentiment classes. |
Named Entity Recognition without Labelled Data: A Weak Supervision Approach (2020.acl-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) performance often degrades when applied to target domains that differ from the texts observed during training. |
| Approach: | They propose a method to learn NER models in the absence of labelled data through weak supervision by using a broad spectrum of labelling functions to automatically annotate texts from the target domain. |
| Outcome: | The proposed approach improves on two English datasets and shows that it improves by 7 percentage points on entity-level F1 scores compared to an out-of-domain neural NER model. |
NERDz: A Preliminary Dataset of Named Entities for Algerian (2022.aacl-short)
Copied to clipboard
| Challenge: | NER is a fundamental task in information extraction and natural language processing. |
| Approach: | They propose to build a manually annotated Algerian vernacular dataset using a recent extension to the Algerian NArabizi Treebank. |
| Outcome: | The proposed dataset is the first of its kind for the Algerian vernacular dialect. |
The interplay between language similarity and script on a novel multi-layer Algerian dialect corpus (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on cross-lingual transfer between languages with similar typology and languages of different scripts. |
| Approach: | They propose to annotate Algerian user-generated comments with parallel annotations . they also investigate the effect of script vs. language similarity in cross-lingual transfer . |
| Outcome: | The proposed model fine-tunes multi-lingual models on Algerian language and scripts . it shows that script vs. language similarity is important for part-of-speech tagging and sentiment analysis . |
NorDiaChange: Diachronic Semantic Change Dataset for Norwegian (2022.lrec-1)
Copied to clipboard
| Challenge: | NorDiaChange is the first dataset of diachronic semantic change on the lexical level for Norwegian. |
| Approach: | They describe a manual annotation process for a new dataset of diachronic semantic change for Norwegian. |
| Outcome: | The proposed dataset covers the time periods related to pre- and post-war events, oil and gas discovery in Norway, and technological developments. |
EDEN: A Dataset for Event Detection in Norwegian News (2024.lrec-main)
Copied to clipboard
Samia Touileb, Jeanett Murstad, Petter Mæhlum, Lubos Steskal, Lilja Charlotte Storset, Huiling You, Lilja Øvrelid
| Challenge: | EDEN is the first dataset annotated with event information at the sentence level for the Norwegian language. |
| Approach: | They propose to annotate Norwegian news text and transcribed speech using ACE event schema. |
| Outcome: | The proposed dataset is the first annotated dataset for Norwegian, with a language-specific annotation process. |