Challenge: Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities.
Approach: They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness .
Outcome: Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds .

Similar Papers

Towards Benchmarking Situational Awareness of Large Language Models:Comprehensive Benchmark, Evaluation and Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Situational awareness is crucial for decision-making, anticipating potential issues, and adapting to dynamic circumstances.
Approach: They propose a benchmark that covers three tiers of situational awareness capabilities . they conduct extensive experiments on advanced LLMs including GPT-4, LLaMA3, Qwen1.5 .
Outcome: The proposed benchmark covers environment perception, situation comprehension and future projection.
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills.
Approach: They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions.
Outcome: The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance.
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities on general text, but their proficiency in specialized scientific domains remains uncharacterized.
Approach: They evaluate the capabilities of large language models in metabolomics research using MetaBench . they found that models perform well on text generation tasks, but cross-database identifier grounding remains challenging .
Outcome: The evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks.
Do Large Language Models Know How Much They Know? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models are highly capable systems, but their capabilities and limitations are unclear.
Approach: They develop a benchmark that challenges LLMs to recall all information they possess on specific topics.
Outcome: The proposed model can recall excessive, insufficient, or the precise amount of information they possess on a given topic, indicating their awareness of how much they know about the given topic.
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have showcased significant improvements in mathematics, but traditional benchmarks like GSM8k offer a unidimensional perspective.
Approach: MathBench is a benchmark that rigorously assesses the mathematical capabilities of large language models.
Outcome: MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
ACEBench: A Comprehensive Evaluation of LLM Tool Usage (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead.
Approach: ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Outcome: ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
CogLM: Tracking Cognitive Development of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have recently shown remarkable abilities across a wide variety of tasks, but few studies have explored the reasons behind the evolutionary relationship among various abilities.
Approach: They construct a benchmark CogLM based on Piaget's Theory of Cognitive Development (PTC) which measures the cognitive levels of Large Language Models (LLMs) using 1,220 questions spanning 10 cognitive abilities crafted by more than 20 human experts.
Outcome: The proposed framework provides a comprehensive testbed for the cognitive levels of LLMs.
Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life .
Approach: They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities.
Outcome: The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities.
SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes.
Approach: They propose a benchmark to assess LLMs' alignment with real-world social attitudes . they find LLM models achieve only 30–40% accuracy when simulating individuals .
Outcome: The proposed benchmark shows that LLMs achieve only 30% accuracy when simulating individuals in complex survey scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations