Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? (2025.findings-acl)
Copied to clipboard
| Challenge: | Current 3D LLMs are evaluated on Q&A or captioning tasks rather than specific downstream tasks like object detection. |
| Approach: | They propose principles for better assessing genuine 3D understanding by explicitly separating 3D abilities from 1D or 2D aspects when evaluating 3D LLMs. |
| Outcome: | The proposed methods are based on the “2D-Cheating” problem in 3D LLM evaluation, suggesting that they are ineffective . |
Similar Papers
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability. |
| Approach: | They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria. |
| Outcome: | The proposed system is based on 11 common aspects with different evaluation criteria. |
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions (2025.emnlp-main)
Copied to clipboard
Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur
| Challenge: | Exact label definitions are considered as clues to disambiguate unclear labels, helping models perform their tasks more effectively. |
| Approach: | They conducted controlled experiments on multiple explanation benchmark datasets and label definition conditions using expert-curated, LLM-generated, perturbed, and swapped definitions. |
| Outcome: | The results suggest that models often default to internal representations, particularly in general tasks, while domain-specific tasks benefit more from explicit definitions. |
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)
Copied to clipboard
Geng Zhang, Yizhou Ying, Sihang Jiang, Jiaqing Liang, Guanglei Yue, Yifei Fu, Hailin Hu, Yanghua Xiao
| Challenge: | Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills. |
| Approach: | They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions. |
| Outcome: | The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance. |
ARC ‘Challenge’ Is Not That Challenging (2025.findings-acl)
Copied to clipboard
| Challenge: | ARC Challenge appears to be more difficult than ARC Easy for modern LLMs due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. |
| Approach: | They propose a setup where multiple choice problems are evaluated and the one with the highest likelihood is compared against the gold standard to determine accuracy. |
| Outcome: | The proposed evaluation setup is more difficult than ARC Easy for modern LLMs because it prevents direct comparison of answer choices rather than inherent complexity. |
CUTE: Measuring LLMs’ Understanding of Their Tokens (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) perform well on a wide variety of tasks, authors say . they lack direct access to characters, which can be difficult to generalize to new languages . |
| Approach: | They propose a benchmark to test the orthographic knowledge of Large Language Models . they find that most LLMs seem to know the spelling of their tokens - yet fail to manipulate text . |
| Outcome: | The proposed benchmark tests the orthographic knowledge of large language models . it finds that most LLMs seem to know the spelling of their tokens, but fail to manipulate text . |
The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators (2025.findings-emnlp)
Copied to clipboard
| Challenge: | LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. |
| Approach: | They revisit meta-evaluations of LLM evaluators under a setting that more closely aligns with practice by examining evaluers’ ability to distinguish test system pairs that are closer in capability. |
| Outcome: | The proposed meta-evaluation setting is significantly different from the use of human evaluations. |
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)
Copied to clipboard
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
| Challenge: | Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution. |
| Approach: | They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. |
| Outcome: | The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage. |
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a bottleneck. |
| Approach: | They propose a visual perception benchmark to test the visual perception of MLLMs. |
| Outcome: | The proposed benchmark examines MLLMs' visual perception abilities with 1758 images and 2612 questions. |
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models are often judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands. |
| Approach: | They propose a diagnostic framework that decomposes benchmark performance into ten cognitively grounded abilities and computes an Ability Impact Score (AIS) AIS quantifies how much each ability contributes to a model’s success on a given benchmark. |
| Outcome: | The proposed framework decomposes performance into ten cognitively grounded abilities and computes an Ability Impact Score (AIS) that quantifies how much each ability contributes to a model’s success on a given benchmark. |
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation (2025.naacl-long)
Copied to clipboard
Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang
| Challenge: | Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases. |
| Approach: | They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process. |
| Outcome: | The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations. |