Papers by Michael Cogswell
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Vision-language models have demonstrated strong efficacy as visual assistants . however, evaluation of their reasoning capabilities requires a costly benchmark . |
| Approach: | They propose a pipeline to measure the reasoning consistency of vision-language models . they propose supervised fine-tuning of VLMs and feedback from LLMs . |
| Outcome: | The proposed framework reduces cost while ensuring the generation of a high-quality dataset. |
BloomVQA: Assessing Hierarchical Multi-modal Comprehension (2024.findings-acl)
Copied to clipboard
Yunye Gong, Robik Shrestha, Jared Claypoole, Michael Cogswell, Arijit Ray, Christopher Kanan, Ajay Divakaran
| Challenge: | Recent advances of machine intelligence solutions have demonstrated tremendous success in a wide range of language and multi-modal tasks over diverse domains. |
| Approach: | They propose a VQA dataset to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. |
| Outcome: | The proposed dataset shows improved accuracy over all comprehension levels and a tendency to bypass visual inputs especially for higher-level tasks. |