Challenge: Current 3D LLMs are evaluated on Q&A or captioning tasks rather than specific downstream tasks like object detection.
Approach: They propose principles for better assessing genuine 3D understanding by explicitly separating 3D abilities from 1D or 2D aspects when evaluating 3D LLMs.
Outcome: The proposed methods are based on the “2D-Cheating” problem in 3D LLM evaluation, suggesting that they are ineffective .

Similar Papers

Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability.
Approach: They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria.
Outcome: The proposed system is based on 11 common aspects with different evaluation criteria.
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions (2025.emnlp-main)

Copied to clipboard

Challenge: Exact label definitions are considered as clues to disambiguate unclear labels, helping models perform their tasks more effectively.
Approach: They conducted controlled experiments on multiple explanation benchmark datasets and label definition conditions using expert-curated, LLM-generated, perturbed, and swapped definitions.
Outcome: The results suggest that models often default to internal representations, particularly in general tasks, while domain-specific tasks benefit more from explicit definitions.
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills.
Approach: They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions.
Outcome: The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance.
ARC ‘Challenge’ Is Not That Challenging (2025.findings-acl)

Copied to clipboard

Challenge: ARC Challenge appears to be more difficult than ARC Easy for modern LLMs due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity.
Approach: They propose a setup where multiple choice problems are evaluated and the one with the highest likelihood is compared against the gold standard to determine accuracy.
Outcome: The proposed evaluation setup is more difficult than ARC Easy for modern LLMs because it prevents direct comparison of answer choices rather than inherent complexity.
CUTE: Measuring LLMs’ Understanding of Their Tokens (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) perform well on a wide variety of tasks, authors say . they lack direct access to characters, which can be difficult to generalize to new languages .
Approach: They propose a benchmark to test the orthographic knowledge of Large Language Models . they find that most LLMs seem to know the spelling of their tokens - yet fail to manipulate text .
Outcome: The proposed benchmark tests the orthographic knowledge of large language models . it finds that most LLMs seem to know the spelling of their tokens, but fail to manipulate text .
The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators (2025.findings-emnlp)

Copied to clipboard

Challenge: LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation.
Approach: They revisit meta-evaluations of LLM evaluators under a setting that more closely aligns with practice by examining evaluers’ ability to distinguish test system pairs that are closer in capability.
Outcome: The proposed meta-evaluation setting is significantly different from the use of human evaluations.
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution.
Approach: They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage.
Outcome: The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage.
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a bottleneck.
Approach: They propose a visual perception benchmark to test the visual perception of MLLMs.
Outcome: The proposed benchmark examines MLLMs' visual perception abilities with 1758 images and 2612 questions.
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models are often judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands.
Approach: They propose a diagnostic framework that decomposes benchmark performance into ten cognitively grounded abilities and computes an Ability Impact Score (AIS) AIS quantifies how much each ability contributes to a model’s success on a given benchmark.
Outcome: The proposed framework decomposes performance into ten cognitively grounded abilities and computes an Ability Impact Score (AIS) that quantifies how much each ability contributes to a model’s success on a given benchmark.
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases.
Approach: They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process.
Outcome: The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations