Papers with Hausa
Thesis Proposal: Self-Adaptive and Epistemic Uncertainty-Guided ASR of Dense Intra-Sentential Code-Switched Speech for African Low-Resource Languages (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing multilingual and pretrained ASR systems improve general recognition accuracy but are weak at switch regions and are sensitive to language imbalance during adaptation. |
| Approach: | They propose a self-adaptive and epistemic uncertainty-guided framework for African low-resource code-switched ASR using Hausa–English and Hausa-Yorùbá as case studies. |
| Outcome: | The proposed framework is based on Hausa–English and Hausa-Yorùbá as case studies. |
Speech Resources in the Tamasheq Language (2022.lrec-1)
Copied to clipboard
Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche, Loïc Barrault, Mickael Rouvier, Yannick Estève
| Challenge: | In this paper, we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger . we share unlabeled audio data in five languages: french, Fulfulde, Hausa, Tamaheq and Zarma . |
| Approach: | They present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. |
| Outcome: | The proposed datasets are used in the IWSLT 2022 low-resource speech translation track . they consist of radio recordings from daily broadcast news in Niger and Mali . |
Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation (2022.lrec-1)
Copied to clipboard
Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Muhammad, Ibrahim Sa’id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello
| Challenge: | Hausa is considered a low resource language in natural language processing due to lack of resources. |
| Approach: | They propose a dataset that contains the description of an image in Hausa and its equivalent in English. |
| Outcome: | The Hausa Visual Genome is the first dataset of its kind . it can be used for Hausa-English machine translation, multi-modal research, image description . |
AFRIDOC-MT: Document-level MT Corpus for African Languages (2025.emnlp-main)
Copied to clipboard
Jesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow
| Challenge: | AFRIDOC-MT is a document-level multi-parallel translation dataset covering five languages . AFRITIC-MT models perform better on sentences than general-purpose LLMs . |
| Approach: | They propose a document-level multi-parallel translation dataset covering English and five African languages. |
| Outcome: | The proposed dataset covers 334 health and 271 information technology news documents . it shows that NLLB-200 achieves the best average performance among standard models . |