Papers by Babak Damavandi

10 papers
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code (2026.eacl-industry)

Copied to clipboard

Challenge: Existing benchmarks do not capture the complexity of structured, step-by-step reasoning essential in physics and related domains.
Approach: They propose a large-scale synthetic benchmark of 15K university-level physics problems . they use structured, step-by-step reasoning and executable Python code to produce the ground-truth solution.
Outcome: The proposed model is based on a set of 15K university-level physics problems with three question types.
SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations (2021.emnlp-main)

Copied to clipboard

Challenge: Existing task-oriented dialog datasets do not situate the dialog in the user’s multimodal context.
Approach: They propose to use a dataset to study multimodal task-oriented dialogs in the shopping domain to situate them in the user’s multimodal context.
Outcome: The proposed dataset includes 11K task-oriented user->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.
VideoMind: Thinking in Steps for Long Video Understanding (2026.eacl-industry)

Copied to clipboard

Challenge: Multimodal Large Language Models struggle with Long Video Understanding due to their limited context window and the distributed nature of salient information across many redundant frames.
Approach: They propose a training framework that mimics a human reasoning process to train Long Video Understanding models.
Outcome: The proposed framework achieves 77.6% performance on Video MME, LongVideo, and MLVU benchmarks while yielding 5% improvement on Llama 4 Scout.
Navigating Connected Memories with a Task-oriented Dialog System (2022.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen an increasing trend in the volume of personal media captured by users thanks to smartphones and smart glasses.
Approach: They propose to use dialogs for connected memories to query media collection . they use a multimodal dialog simulator and manual paraphrasing to obtain natural language utterances.
Outcome: The proposed dataset contains 11.5k userassistant dialogs grounded in simulated personal memory graphs.
SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams (2023.acl-long)

Copied to clipboard

Challenge: Existing models lack a large-scale benchmark to capture user–assistant interactions . et al., 2022: 145-160.
Approach: They propose a video-grounded task-oriented dialog dataset that captures real-world AI-assisted user scenarios in VR.
Outcome: The proposed dataset captures real-world AI-assisted user scenarios in VR.
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM (2024.findings-emnlp)

Copied to clipboard

Challenge: Vision-extended LLMs have made significant strides in VQA, but they still encounter significant difficulties in handling queries involving long-tail entities.
Approach: They propose a benchmark to test models' ability to identify entities and provide detailed, entity-specific knowledge by combining 10 images and 10 knowledge-intensive QA pairs.
Outcome: The proposed model outperforms existing methods on the SnapNTell dataset, achieving a 66.5% improvement in the BELURT score.
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Evaluating large vision-language models has focused on final-answer correctness, but this metric is often insufficient and misleading.
Approach: They propose a framework that decomposes complex multimodal tasks into Auxiliary Reasoning Sets (ARS) ARS decomposition reveals how consistently a model reasons across sub-questions with structured dependencies.
Outcome: a new framework improves diagnostic evaluation of large vision-language models . it decomposes complex multimodal tasks into auxiliary reasoning sets with structured dependencies . the framework pinpoints reasoning failures and exposes errors overlooked by standard evaluation .
IMU2CLIP: Language-grounded Motion Sensor Translation with Multimodal Contrastive Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to align motion sensors with text and video are limited in their scale and limited in the use of IMU models.
Approach: They propose to project IMU motion sensor recordings into the joint representation space of Contrastive Language-Image Pre-training (CLIP) they introduce several new IMU-based Wearable AI applications such as motion-based media search, or an LM-based multimodal reasoning with motion sensor data.
Outcome: The proposed approach significantly improves downstream performance when fine-tuned for each application, demonstrating its universal usage as a new pre-trained resource.
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in conversational AI have been substantial, but developing real-time tasks guidance systems remains a challenge.
Approach: They propose a data curation pipeline that synthesizes dialogues from annotated egocentric videos and a suite of automatic evaluation metrics that validated through extensive human studies.
Outcome: The proposed framework synthesizes dialogues from annotated egocentric videos and validates them through extensive human studies.
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model (2024.emnlp-industry)

Copied to clipboard

Challenge: Prior work on LLMs focused on models that combine text and one other modality, such as image encoders or proprietary models that are not open sourced.
Approach: They propose a unified model that reasons over diverse input modality signals and generates textual responses.
Outcome: The proposed model performs better on multimodal tasks than industry-leading models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations