CogEvolve: A Multimodal Benchmark for Evaluating Relational Reasoning in Semantic Extension (2026.acl-long)
Copied to clipboard
| Challenge: | a gap exists between human embodied logic and machine statistical learning . authors: models internalize statistical patterns or mimic static recognition . |
| Approach: | They propose a cognitive linguistic benchmark to test whether large language models internalize statistical logic or not . they find that models function as "Super-Associators" expert at static recognition yet fail at causal reasoning . |
| Outcome: | The proposed model fails at causal reasoning and has a high-fidelity concept representation but lacks transformational operators essential for true relational understanding. |
Similar Papers
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal large language models (MLLMs) have demonstrated compelling visual understanding in recent years. |
| Approach: | They propose a multimodal large language model with eight visuo-cognitive tasks inspired by classic human intelligence tests organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
| Outcome: | The proposed frameworks are based on eight visuo-cognitive tasks inspired by human intelligence tests and organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark (2026.acl-long)
Copied to clipboard
Kai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu
| Challenge: | Existing evaluations treat visual understanding and generation in isolation or overlook tasks that inherently couple them. |
| Approach: | They propose a benchmark that examines the bidirectional synergy between generation and understanding across eight reasoning-centric domains. |
| Outcome: | The proposed model systematically unfolds the bidirectional synergy between generation and understanding across eight reasoning-centric domains. |
MiQA: A Benchmark for Inference on Metaphorical Questions (2022.aacl-short)
Copied to clipboard
| Challenge: | a benchmark is proposed to assess the capability of large language models to reason with conventional metaphors. |
| Approach: | They propose to assess the capability of large language models to reason with conventional metaphors. |
| Outcome: | The proposed benchmark compares pre-trained models on binary-choice tasks with human models . the results show that human models perform better on the largest model, compared to small models based on the same task . |
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for Multimodal Large Language Models (MLLMs) have been lacking due to the rich nature of social interaction. |
| Approach: | They propose a video benchmark to evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction. |
| Outcome: | The proposed benchmarks evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction tasks. |
CHARPEVAL: Benchmarking Large Language Models’ Contextual Reasoning in Knowledge-Grounded Dialogue (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks that evaluate the ability of Large Language Models (LLMs) to perform contextualized reasoning in knowledge-grounded dialogue scenarios are lacking. |
| Approach: | They propose a benchmark to evaluate the ability of Large Language Models to perform contextualized reasoning in knowledge-grounded dialogue scenarios. |
| Outcome: | The proposed benchmark shows that open-weight LLMs are ineffective at reasoning over discontinuous chunks of text across the input. |
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts (2025.findings-acl)
Copied to clipboard
Negar Foroutan, Angelika Romanou, Matin Ansaripour, Julian Martin Eisenschlos, Karl Aberer, Rémi Lebret
| Challenge: | Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). |
| Approach: | They propose a benchmark for evaluating cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages . they evaluate 12 vision-language models that achieve 70% accuracy when provided with direct context . |
| Outcome: | The proposed benchmark evaluates models with high accuracy over tables and charts extracted from 4,000 Wikipedia pages . proprietary models achieve 70% accuracy when provided with direct context, but open-source models perform worse when retrieval from long documents is required. |
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation (2025.naacl-long)
Copied to clipboard
Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang
| Challenge: | Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases. |
| Approach: | They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process. |
| Outcome: | The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations. |
Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on metaphor processing have focused on single datasets and specific task settings, often using artificially constructed data through lexical replacement. |
| Approach: | They propose to evaluate the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations. |
| Outcome: | The proposed frameworks are more realistic and efficient than current models and are more efficient than existing models. |
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models (2024.findings-emnlp)
Copied to clipboard
Yaojia Lv, Haojie Pan, Zekun Wang, Jiafeng Liang, Yuanxing Liu, Ruiji Fu, Ming Liu, Zhongyuan Wang, Bing Qin
| Challenge: | Recent advances in large language models (LLMs) focus on replicating human cognition in specific contexts, overlooking the inherently dynamic nature of cognition. |
| Approach: | They propose a task to assess cognitive dynamics of large language models (LLMs) they introduce a benchmark and two evaluation metrics to validate the benchmark and evaluate it through participant surveys. |
| Outcome: | The proposed task overcomes the limitations of existing methods and is available for download. |
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding (2026.findings-eacl)
Copied to clipboard
| Challenge: | figurative language is essential for expressing intent, emotion, and perspective . figural language is often dependent on Styles Reasoning, causing incongruities between expressions . |
| Approach: | They propose a framework that induces reasoning capabilities to compact vision–language models . figurative language is essential in expressing intent, emotion, and perspective . |
| Outcome: | The proposed framework can interpret multimodal figurative language, provide transparent reasoning traces, and generalize across multiple figurativ styles. |