Papers by Colin Leong

5 papers
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries.
Approach: They propose a large and highly multilingual dataset for sign language translation: JWSign.
Outcome: The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers.
Phone-ing it in: Towards Flexible Multi-Modal Language Model Training by Phonetic Representations of Data (2022.acl-long)

Copied to clipboard

Challenge: Pre-trained language models are increasingly applied in ways that are agnostic to targeted downstream tasks.
Approach: They propose a multi-modal approach to train language models using whatever text and/or audio data might be available in a language.
Outcome: The proposed approach improves on pre-trained models on Swahili and Kinyarwanda data, with an improvement of up to 6% over models that are trained from scratch.
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families.
Approach: They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition.
Outcome: The Bloom Library datasets cover 363 languages across 32 language families.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations