Papers by Sahiti Yerramilli
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Leveraging 1.46 million Mapillary street-level images, GeoChain pairs each image with a 21-step chain-of-thought (CoT) question sequence (over 30 million Q&A pairs). |
| Approach: | They propose a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs) they pair each image with a 21-step chain-of-thought (CoT) question sequence (over 30 million Q&A pairs) |
| Outcome: | The proposed benchmark pairs 1.46 million images with a 21-step chain-of-thought (CoT) question sequence (over 30 million Q&A pairs) |