Papers by Joshua Nemecek
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks (2022.emnlp-main)
Copied to clipboard
| Challenge: | In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. |
| Approach: | They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. |
| Outcome: | The Bloom Library datasets cover 363 languages across 32 language families. |