Papers by Colin Leong
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries. |
| Approach: | They propose a large and highly multilingual dataset for sign language translation: JWSign. |
| Outcome: | The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022.tacl-1)
Copied to clipboard
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofetoluwa Adeyemi
| Challenge: | Lower-resource corpora have systematic issues, including mislabeled or nonstandard/ambiguous language codes. |
| Approach: | They manually audit the quality of 205 language-specific corpora released with five major public datasets. |
| Outcome: | The results show that lower-resource corpora have systematic issues even for non-proficient speakers. |
Phone-ing it in: Towards Flexible Multi-Modal Language Model Training by Phonetic Representations of Data (2022.acl-long)
Copied to clipboard
| Challenge: | Pre-trained language models are increasingly applied in ways that are agnostic to targeted downstream tasks. |
| Approach: | They propose a multi-modal approach to train language models using whatever text and/or audio data might be available in a language. |
| Outcome: | The proposed approach improves on pre-trained models on Swahili and Kinyarwanda data, with an improvement of up to 6% over models that are trained from scratch. |
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks (2022.emnlp-main)
Copied to clipboard
| Challenge: | In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. |
| Approach: | They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. |
| Outcome: | The Bloom Library datasets cover 363 languages across 32 language families. |
A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation (2022.naacl-main)
Copied to clipboard
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, Sam Manthalu
| Challenge: | Low-resource languages are left out of large-scale pretraining datasets . authors explore how to leverage existing pre-trained models to create low-resourced translation systems for 16 African languages. |
| Approach: | They investigate how large-scale pre-trained models can be used to create low-resource translation systems for 16 African languages. |
| Outcome: | The proposed models can translate between hundreds of languages even though there is little parallel data available for training. |