Papers by Weidong Cai

4 papers
Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings.
Approach: They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs.
Outcome: The proposed benchmark leverages siamese images and text pairs to challenge MLLMs.
Enhancing Advanced Visual Reasoning Ability of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability.
Approach: They propose a novel multi-modal in-context learning methodology to enhance LLMs’ contextual understanding and reasoning.
Outcome: The proposed model achieves SOTA performance among all visual reasoning tasks and achieves a 'higher level of accuracy' than previous models.
Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding (2026.acl-long)

Copied to clipboard

Challenge: Existing vision-language models suffer from visual hallucination, where the generated responses contain inaccuracies that are not grounded in the visual input.
Approach: They propose a visual decoding method that iteratively isolates relevant visual tokens based on attention and magnifies the corresponding regions.
Outcome: The proposed method reduces language biases and amplifies weights of visual embedding during decoding, while still preserving strong reasoning capabilities.
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)

Copied to clipboard

Challenge: Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training.
Approach: They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.
Outcome: The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations