Papers by Shao-Yen Tseng
ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning (2023.acl-long)
Copied to clipboard
Xiao Xu, Bei Li, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan
| Challenge: | Two-Tower Vision-Language models suffer from ineffective layer-by-layer utilization of uni-modal representations and cannot flexibly exploit different levels of unil-modal knowledge. |
| Approach: | They propose a model architecture that gathers and combines the insights of pre-trained uni-modal experts at different levels to facilitate more comprehensive cross-modal alignment and fusion. |
| Outcome: | The proposed model outperforms baselines with and without Vision-Language Pre-training (VLP) with 4M VLP data. |
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing vision-and-language pretraining approaches rely on external object detectors to encode images in a multi-modal transformer framework. |
| Approach: | They propose an object-aware end-to-end VLP framework which feeds image grid features from CNNs into the Transformer and learns the multi-modal representations jointly. |
| Outcome: | The proposed framework achieves competitive or superior performances on vision-language tasks. |
Why do LLaVA Vision-Language Models Reply to Images in English? (2024.findings-emnlp)
Copied to clipboard
Musashi Hinck, Carolin Holtermann, Matthew Olson, Florian Schneider, Sungduk Yu, Anahita Bhiwandiwalla, Anne Lauscher, Shao-Yen Tseng, Vasudev Lal
| Challenge: | Including an image in a multimodal query significantly increases the likelihood of the model returning an English response regardless of the language of the query. |
| Approach: | They propose a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models’ internal representations of image and text inputs. |
| Outcome: | The proposed approach reduces the multilingual error by switching the language backbone for a bilingual language model. |