Challenge: Large Language Models (LLMs) have recently shown remarkable abilities across a wide variety of tasks, but few studies have explored the reasons behind the evolutionary relationship among various abilities.
Approach: They construct a benchmark CogLM based on Piaget's Theory of Cognitive Development (PTC) which measures the cognitive levels of Large Language Models (LLMs) using 1,220 questions spanning 10 cognitive abilities crafted by more than 20 human experts.
Outcome: The proposed framework provides a comprehensive testbed for the cognitive levels of LLMs.

Similar Papers

Development of Cognitive Intelligence in Pre-trained Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies show evidence for emergent cognitive abilities in Large Pre-trained Language Models (PLMs). Prior research into emergental cognitive abilities of PLMs has been path-independent to model training.
Approach: They use four task categories to examine the alignment of ten popular families of PLMs and evaluate their performance to the developmental trajectories of children's thinking.
Outcome: The results show that the models are more aligned to children's thinking than previous studies.
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) focus on replicating human cognition in specific contexts, overlooking the inherently dynamic nature of cognition.
Approach: They propose a task to assess cognitive dynamics of large language models (LLMs) they introduce a benchmark and two evaluation metrics to validate the benchmark and evaluate it through participant surveys.
Outcome: The proposed task overcomes the limitations of existing methods and is available for download.
Working Memory Identifies Reasoning Limits in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we examine the limitations of their cognitive capabilities and their working memory.
Approach: They examine the limitations of large language models from a scaling perspective . they also assess various prompting strategies, revealing their diverse impacts on LLM performance.
Outcome: The proposed models perform poorly on n-back tasks and on prompting strategies.
Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on LLMs evaluation with exams are lacking in cognitive research on their overall knowledge structure.
Approach: They conduct an evaluation using a human test dataset based on Bloom Taxonomy to reveal the knowledge structures of Large Language Models and gain insights of their cognitive capabilities.
Outcome: The proposed model can pass AP, SAT, and Leetcode exams, but lacks the cognitive power to perform on human exams.
CogBench: Benchmarking Cognitive Alignment of Large Language Models in Educational Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) possess strong capabilities in language understanding and generation, as well as remarkable problem-solving abilities.
Approach: They propose a benchmark to assess the cognitive alignment capabilities of large language models in educational QA.
Outcome: The proposed evaluation benchmark assesses the cognitive alignment capabilities of large language models in educational QA.
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.
Approach: They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning.
Outcome: The proposed model can generate grammatically correct, fluent text without requiring human intervention.
From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets focus on low-level cognitive tasks while providing limited coverage of higher-level reasoning skills.
Approach: They analyze the cognitive depth of popular LLM benchmarks using Bloom’s Taxonomy to evaluate both the cognitive and knowledge dimensions.
Outcome: The results show that incorporating higher-level cognitive instructions into the current instruction fine-tuning process improves model performance.
Evaluating the Deductive Competence of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance .
Approach: They investigate whether several large language models can solve a deductive reasoning problem in their conventional form.
Outcome: The proposed models can solve a classic type of deductive reasoning problem in their conventional form.
LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation approaches for Large Language Models lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks.
Approach: They propose a hierarchical approach to evaluation of Large Language Models that leverages Bloom’s Taxonomy to identify how well they cover the levels of Bloom’ s taxonomies.
Outcome: The proposed evaluation frameworks cover the Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills, on the most widely used benchmarks.
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning.
Approach: They introduce a dataset grounded in the pragmatic concept of alternatives to evaluate whether large language models can accurately infer nuanced speaker intentions.
Outcome: The proposed model can infer nuanced speaker intentions by inferring the speaker’s intended meaning and explaining when and why a speaker would choose one utterance over its alternative.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations