CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language (2026.acl-long)
Copied to clipboard
| Challenge: | CNSL-bench is the first comprehensive Chinese National Sign Language benchmark . current MLLMs are inferior to human performance, despite advances in multimodal modeling . |
| Approach: | They propose a Chinese National Sign Language benchmark to evaluate multimodal large language models in sign language understanding. |
| Outcome: | The proposed benchmark evaluates 21 open-source and proprietary MLLMs . results show that current models are inferior to human performance . |
Similar Papers
MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical Studies (2025.acl-long)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have advanced visual and language understanding, but their potential in Chinese Classical Studies (CCS) remains underexplored due to the lack of specialized benchmarks. |
| Approach: | They propose to develop a multimodal benchmark specifically designed for Chinese Classical Studies across multiple subdomains to bridge this gap. |
| Outcome: | The proposed benchmark spans seven core subdomains with a total of 45 meticulously designed tasks. |
Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale Recognition (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on a single type of quantity or a specific format, lacking a comprehensive evaluation of scale recognition capabilities. |
| Approach: | They propose a visual scale recognition benchmark built using images from COCO, Open Images, and Flickr to evaluate scale recognition capabilities of multimodal large language models. |
| Outcome: | The proposed model achieves 42.60% accuracy, lower than the 97.40% of humans. |
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts. |
| Approach: | They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users. |
| Outcome: | The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm. |
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)
Copied to clipboard
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen, Junying Chen, Zhihong Chen, Wenya Xie, Shuo Yan, ChenghaoZhu ChenghaoZhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang
| Challenge: | Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences. |
| Approach: | They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge. |
| Outcome: | The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria. |
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to reflect real-world communication needs and are limited in their coverage. |
| Approach: | They present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. |
| Outcome: | The proposed index covers 120 resources across 35 sign languages. |
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain (2024.findings-acl)
Copied to clipboard
Liang Chen, Yichi Zhang, Shuhuai Ren, Haozhe Zhao, Zefan Cai, Yuchi Wang, Peiyi Wang, Xiangdi Meng, Tianyu Liu, Baobao Chang
| Challenge: | a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models. |
| Approach: | They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models . |
| Outcome: | The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions . |
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)
Copied to clipboard
Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang, Yongqian Wen, Yuanxing Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Shihao Li, Yanghai Wang, Tianhao Peng, Jiaheng Liu
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
Can MLLMs Understand the Deep Implication Behind Chinese Images? (2025.acl-long)
Copied to clipboard
Chenhao Zhang, Xi Feng, Yuelin Bai, Xeron Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Jiaheng Liu, Xingwei Qu, Yifei Zhang, Qixuan Zhao, Yiming Liang, Ziqiang Liu, Feiteng Fang, Min Yang, Wenhao Huang, Chenghua Lin, Ge Zhang, Shiwen Ni
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Thesis Proposal: Multimodal Benchmark for Music Understanding in Large Language Models (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing music-focused benchmarks are fragmented, largely single-modality, Western-centric . existing methods for evaluating MLLMs are lacking reproducibility and reliability . |
| Approach: | They propose to develop a musically multimodal benchmark that will integrate music into the benchmark. |
| Outcome: | The proposed benchmark will integrate culturally diverse musical material beyond the dominant Western canon. |
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on basic abilities using nonverbal methods, such as yes-no and multiple-choice questions. |
| Approach: | They propose a benchmark that provides more nuanced evaluations of alignment capabilities for large Vision-Language Models (VLMs) they use a rule-calibrated evaluator that exceeds GPT-4's evaluation ability and a “alignment score” to assess the robustness and stability of models across diverse prompts. |
| Outcome: | The proposed benchmark covers 13 tasks across three categories and includes both single-turn and multi-turn dialogue scenarios. |