Papers by Lisen Dai
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning (2026.acl-long)
Copied to clipboard
Tianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang, Lisen Dai, Chen Zhao, Minhao Cheng, Arman Cohan, Xiangliang Zhang
| Challenge: | Multimodal mathematical Reasoning (MMR) has attracted increasing attention for its ability to solve mathematical problems involving both textual and visual modalities. |
| Approach: | They review the theoretical frameworks of multimodal reasoning and examine the challenges they face in visual math tasks. |
| Outcome: | The proposed models can solve problems involving both textual and visual modalities. |
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP (2025.acl-long)
Copied to clipboard
| Challenge: | MU has gained significant attention as a means to remove the influence of specific data from a trained model without requiring full retraining. |
| Approach: | They propose a novel approach that disentangles and selectively forgets both visual and textual associations, ensuring that unlearning does not compromise model performance. |
| Outcome: | Experiments on CIFAR-100, Flickr30K, and Conceptual 12M show that CLIPErase effectively removes designated associations from multimodal samples in downstream tasks while preserving model performance on retain set. |
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing AVQA methods often fail to link sound-producing objects in the video with the audio-visual information. |
| Approach: | They introduce a source-aware semantic representation network for AVQA . they use source-wise learnable tokens to capture and align audio-visual elements with the question . |
| Outcome: | The proposed model outperforms state-of-the-art models on the Music-AVQA and AVQA-Yang datasets. |