Papers by Hailu Balcha
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens (2026.findings-acl)
Copied to clipboard
Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha, Debora Taye Tesfaye, Elbethel Daniel Zewdie, Ikram Behiru Nesiru, Jitu Ewnetu Hailu, Senait Mengesha Yayo
| Challenge: | afan oromo, amharic, and tigrinya are low-resourced languages . they are used for training, benchmarks, news, health, and sports . afono o'mara: quantity does not guarantee quality of MT datasets . |
| Approach: | They investigate the quality of machine translation datasets for three low-resourced languages . they found a large skew towards the male gender in the datasets . |
| Outcome: | The results show that training data has large representation of political and religious text, but benchmark datasets focus on news, health, and sports. |
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)
Copied to clipboard
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur
| Challenge: | Africa has the highest linguistic diversity among all continents. |
| Approach: | They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges . |
| Outcome: | The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task . |