AwarenessBench: Assessing Cognitive Capabilities of Language Models (2026.acl-long)
Copied to clipboard
Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen, Haoyuan Shi, Yu Wang, Wei Xu
| Challenge: | Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities. |
| Approach: | They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness . |
| Outcome: | Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds . |
Similar Papers
Towards Benchmarking Situational Awareness of Large Language Models:Comprehensive Benchmark, Evaluation and Analysis (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Situational awareness is crucial for decision-making, anticipating potential issues, and adapting to dynamic circumstances. |
| Approach: | They propose a benchmark that covers three tiers of situational awareness capabilities . they conduct extensive experiments on advanced LLMs including GPT-4, LLaMA3, Qwen1.5 . |
| Outcome: | The proposed benchmark covers environment perception, situation comprehension and future projection. |
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)
Copied to clipboard
Geng Zhang, Yizhou Ying, Sihang Jiang, Jiaqing Liang, Guanglei Yue, Yifei Fu, Hailin Hu, Yanghua Xiao
| Challenge: | Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills. |
| Approach: | They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions. |
| Outcome: | The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance. |
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics (2026.acl-long)
Copied to clipboard
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May Dongmei Wang
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities on general text, but their proficiency in specialized scientific domains remains uncharacterized. |
| Approach: | They evaluate the capabilities of large language models in metabolomics research using MetaBench . they found that models perform well on text generation tasks, but cross-database identifier grounding remains challenging . |
| Outcome: | The evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks. |
Do Large Language Models Know How Much They Know? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models are highly capable systems, but their capabilities and limitations are unclear. |
| Approach: | They develop a benchmark that challenges LLMs to recall all information they possess on specific topics. |
| Outcome: | The proposed model can recall excessive, insufficient, or the precise amount of information they possess on a given topic, indicating their awareness of how much they know about the given topic. |
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark (2024.findings-acl)
Copied to clipboard
Hongwei Liu, Zilong Zheng, Yuxuan Qiao, Haodong Duan, Zhiwei Fei, Fengzhe Zhou, Wenwei Zhang, Songyang Zhang, Dahua Lin, Kai Chen
| Challenge: | Recent advances in large language models have showcased significant improvements in mathematics, but traditional benchmarks like GSM8k offer a unidimensional perspective. |
| Approach: | MathBench is a benchmark that rigorously assesses the mathematical capabilities of large language models. |
| Outcome: | MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills. |
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
ACEBench: A Comprehensive Evaluation of LLM Tool Usage (2025.findings-emnlp)
Copied to clipboard
Chen Chen, Xinlong Hao, Weiwen Liu, Xu Huang, Xingshan Zeng, Shuai Yu, Dexun Li, Yuefeng Huang, Xiangcheng Liu, Wang Xinzhi, Wu Liu
| Challenge: | Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead. |
| Approach: | ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. |
| Outcome: | ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. |
CogLM: Tracking Cognitive Development of Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have recently shown remarkable abilities across a wide variety of tasks, but few studies have explored the reasons behind the evolutionary relationship among various abilities. |
| Approach: | They construct a benchmark CogLM based on Piaget's Theory of Cognitive Development (PTC) which measures the cognitive levels of Large Language Models (LLMs) using 1,220 questions spanning 10 cognitive abilities crafted by more than 20 human experts. |
| Outcome: | The proposed framework provides a comprehensive testbed for the cognitive levels of LLMs. |
Theory of Mind in Large Language Models: Assessment and Enhancement (2025.acl-long)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) is a cornerstone of human social intelligence . Large Language Models (LLMs) are increasingly integrated into daily life . |
| Approach: | They analyze evaluation benchmarks and enhancement strategies to evaluate LLMs' ToM capabilities. |
| Outcome: | The proposed and widely used story-based benchmarks and enhancement strategies are used to evaluate LLMs' ToM capabilities. |
SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. |
| Approach: | They propose a benchmark to assess LLMs' alignment with real-world social attitudes . they find LLM models achieve only 30–40% accuracy when simulating individuals . |
| Outcome: | The proposed benchmark shows that LLMs achieve only 30% accuracy when simulating individuals in complex survey scenarios. |