Papers by Babak Damavandi
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing benchmarks do not capture the complexity of structured, step-by-step reasoning essential in physics and related domains. |
| Approach: | They propose a large-scale synthetic benchmark of 15K university-level physics problems . they use structured, step-by-step reasoning and executable Python code to produce the ground-truth solution. |
| Outcome: | The proposed model is based on a set of 15K university-level physics problems with three question types. |
SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing task-oriented dialog datasets do not situate the dialog in the user’s multimodal context. |
| Approach: | They propose to use a dataset to study multimodal task-oriented dialogs in the shopping domain to situate them in the user’s multimodal context. |
| Outcome: | The proposed dataset includes 11K task-oriented user->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes. |
VideoMind: Thinking in Steps for Long Video Understanding (2026.eacl-industry)
Copied to clipboard
Shubhang Bhatnagar, Renxiong Wang, Kapil Krishnakumar, Adel Ahmadyan, Zhaojiang Lin, Lambert Mathias, Xin Luna Dong, Babak Damavandi, Narendra Ahuja, Seungwhan Moon
| Challenge: | Multimodal Large Language Models struggle with Long Video Understanding due to their limited context window and the distributed nature of salient information across many redundant frames. |
| Approach: | They propose a training framework that mimics a human reasoning process to train Long Video Understanding models. |
| Outcome: | The proposed framework achieves 77.6% performance on Video MME, LongVideo, and MLVU benchmarks while yielding 5% improvement on Llama 4 Scout. |
Navigating Connected Memories with a Task-oriented Dialog System (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent years have seen an increasing trend in the volume of personal media captured by users thanks to smartphones and smart glasses. |
| Approach: | They propose to use dialogs for connected memories to query media collection . they use a multimodal dialog simulator and manual paraphrasing to obtain natural language utterances. |
| Outcome: | The proposed dataset contains 11.5k userassistant dialogs grounded in simulated personal memory graphs. |
SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams (2023.acl-long)
Copied to clipboard
Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodriguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon
| Challenge: | Existing models lack a large-scale benchmark to capture user–assistant interactions . et al., 2022: 145-160. |
| Approach: | They propose a video-grounded task-oriented dialog dataset that captures real-world AI-assisted user scenarios in VR. |
| Outcome: | The proposed dataset captures real-world AI-assisted user scenarios in VR. |
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM (2024.findings-emnlp)
Copied to clipboard
Jielin Qiu, Andrea Madotto, Zhaojiang Lin, Paul Crook, Yifan Xu, Babak Damavandi, Xin Dong, Christos Faloutsos, Lei Li, Seungwhan Moon
| Challenge: | Vision-extended LLMs have made significant strides in VQA, but they still encounter significant difficulties in handling queries involving long-tail entities. |
| Approach: | They propose a benchmark to test models' ability to identify entities and provide detailed, entity-specific knowledge by combining 10 images and 10 knowledge-intensive QA pairs. |
| Outcome: | The proposed model outperforms existing methods on the SnapNTell dataset, achieving a 66.5% improvement in the BELURT score. |
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Evaluating large vision-language models has focused on final-answer correctness, but this metric is often insufficient and misleading. |
| Approach: | They propose a framework that decomposes complex multimodal tasks into Auxiliary Reasoning Sets (ARS) ARS decomposition reveals how consistently a model reasons across sub-questions with structured dependencies. |
| Outcome: | a new framework improves diagnostic evaluation of large vision-language models . it decomposes complex multimodal tasks into auxiliary reasoning sets with structured dependencies . the framework pinpoints reasoning failures and exposes errors overlooked by standard evaluation . |
IMU2CLIP: Language-grounded Motion Sensor Translation with Multimodal Contrastive Learning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to align motion sensors with text and video are limited in their scale and limited in the use of IMU models. |
| Approach: | They propose to project IMU motion sensor recordings into the joint representation space of Contrastive Language-Image Pre-training (CLIP) they introduce several new IMU-based Wearable AI applications such as motion-based media search, or an LM-based multimodal reasoning with motion sensor data. |
| Outcome: | The proposed approach significantly improves downstream performance when fine-tuned for each application, demonstrating its universal usage as a new pre-trained resource. |
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos (2025.emnlp-main)
Copied to clipboard
Yichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, Seungwhan Moon
| Challenge: | Recent advances in conversational AI have been substantial, but developing real-time tasks guidance systems remains a challenge. |
| Approach: | They propose a data curation pipeline that synthesizes dialogues from annotated egocentric videos and a suite of automatic evaluation metrics that validated through extensive human studies. |
| Outcome: | The proposed framework synthesizes dialogues from annotated egocentric videos and validates them through extensive human studies. |
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model (2024.emnlp-industry)
Copied to clipboard
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Tushar Nagarajan, Matt Smith, Shashank Jain, Chun-Fu Yeh, Prakash Murugesan, Peyman Heidari, Yue Liu, Kavya Srinet, Babak Damavandi, Anuj Kumar
| Challenge: | Prior work on LLMs focused on models that combine text and one other modality, such as image encoders or proprietary models that are not open sourced. |
| Approach: | They propose a unified model that reasons over diverse input modality signals and generates textual responses. |
| Outcome: | The proposed model performs better on multimodal tasks than industry-leading models . |