Papers by Aashi Jain
MURAL: Multimodal, Multitask Representations Across Languages (2021.findings-emnlp)
Copied to clipboard
Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, Jason Baldridge
| Challenge: | Image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. |
| Approach: | They propose a dual encoder that integrates image-text matching and translation pairs to solve two tasks by learning from billions of pairs. |
| Outcome: | The proposed encoder outperforms ALIGN's cross-modal retrieval performance on well-resourced languages and significantly improves on under-resource languages. |