Papers by Souhir Gahbiche
Speech Resources in the Tamasheq Language (2022.lrec-1)
Copied to clipboard
Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche, Loïc Barrault, Mickael Rouvier, Yannick Estève
| Challenge: | In this paper, we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger . we share unlabeled audio data in five languages: french, Fulfulde, Hausa, Tamaheq and Zarma . |
| Approach: | They present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. |
| Outcome: | The proposed datasets are used in the IWSLT 2022 low-resource speech translation track . they consist of radio recordings from daily broadcast news in Niger and Mali . |
K-pop and fake facts: from texts to smart alerting for maritime security (2023.acl-industry)
Copied to clipboard
Maxime Prieur, Souhir Gahbiche, Guillaume Gadek, Sylvain Gatepaille, Kilian Vasnier, Valerian Justine
| Challenge: | Maritime security requires full-time monitoring of the situation based on technical data but also from OSINT-like inputs. |
| Approach: | They propose a system that extracts data from sensors and texts to feed a Knowledge Base . the system can be used to detect malicious actors using AIS and pseudo-newspapers . |
| Outcome: | The proposed system ingests data from sensors and texts and feeds a Knowledge Base . it performs coherence checks between extracted facts and the extracted data . |
Arabizi Language Models for Sentiment Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Arabizi is a written form of spoken Arabic, relying on Latin characters and digits. |
| Approach: | They propose to use Arabizi as a written form of spoken Arabic in online social networks . they use a corpus of 7.7M tweets written in Arabizi and a subset of SALAD to train a model in Arabic . |
| Outcome: | The proposed model outperforms state-of-the-art models on sentiment analysis task using arabizi . the proposed model is based on a corpus of 7.7M tweets written in arabizi and a subset of LAD manually annotated for sentiment analysis. |