Papers by Saehoon Kim
Efficient Multilingual Multi-modal Pre-training through Triple Contrastive Loss (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches to learn visual and textual representations from web-scale image-text pairs are limited due to labeling cost and limited scalability. |
| Approach: | They propose to use web-scale image-text pairs to learn visual and textual representations in the shared space. |
| Outcome: | The proposed enhancement scheme improves multilingual vision-and-language tasks by minimizing a triplet contrastive loss on images and two different language texts with the same meaning. |