Papers by Abhanshu Sharma
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs (2024.findings-naacl)
Copied to clipboard
Victor Carbune, Hassan Mansoor, Fangyu Liu, Rahul Aralikatte, Gilles Baechler, Jindong Chen, Abhanshu Sharma
| Challenge: | Visual language models (VLMs) are achieving increasingly strong performance on multimodal tasks. |
| Approach: | They propose to transfer reasoning capabilities from large-language models to VLMs by constructing a 20x larger dataset and a larger dataset to improve general reasoning capabilities. |
| Outcome: | The proposed model outperforms larger models without an upstream OCR system while keeping inference time constant. |
Towards Better Semantic Understanding of Mobile Interfaces (2022.coling-1)
Copied to clipboard
Srinivas Sunkara, Maria Wang, Lijuan Liu, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, James W. W. Stout
| Challenge: | a dataset of 500k unique annotations is released to improve mobile accessibility and automation capabilities. |
| Approach: | They propose to use an annotation dataset to improve the accessibility of mobile UIs . they use images and view hierarchies to augment annotations for icons and their semantics - and use multimodal inputs to build models. |
| Outcome: | The proposed dataset shows that it can be used to improve UIs and categories on unseen apps. |