Challenge: Recent research on fluid intelligence assessments has highlighted significant deficiencies in LLMs’ abilities.
Approach: They analyze the challenges LLMs face in demonstrating fluid intelligence through controlled experiments using the most representative ARC task as an example.
Outcome: The proposed model shows that it lacks the ability to combine skill composition and abstract input formats and lacks left-to-right decoding.

Similar Papers

ARC ‘Challenge’ Is Not That Challenging (2025.findings-acl)

Copied to clipboard

Challenge: ARC Challenge appears to be more difficult than ARC Easy for modern LLMs due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity.
Approach: They propose a setup where multiple choice problems are evaluated and the one with the highest likelihood is compared against the gold standard to determine accuracy.
Outcome: The proposed evaluation setup is more difficult than ARC Easy for modern LLMs because it prevents direct comparison of answer choices rather than inherent complexity.
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability.
Approach: They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria.
Outcome: The proposed system is based on 11 common aspects with different evaluation criteria.
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in comprehending and analyzing lengthy sequential inputs.
Approach: They propose to implement ad-hoc solutions that enhance LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to address this limitation, they propose to use three datasets and two tasks to analyze news categorization and sentence analysis to evaluate their models.
Outcome: The proposed solutions significantly improve LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to 93% and 50%, respectively.
LLM-driven Instruction Following: Progresses and Concerns (2023.emnlp-tutorial)

Copied to clipboard

Challenge: a tutorial on task instruction is aimed at researchers and practitioners interested in NLP generalization . labeled examples are unlikely to be available in large numbers or do not exist .
Approach: This tutorial will examine the progress of natural language processing (NLP) using labeled examples. authors propose that task instructions act as a novel resource for supervision.
Outcome: This tutorial aims to answer questions about instruction-driven NLP . it focuses on the use of task instructions in a low-shot scenario .
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems.
Approach: They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM.
Outcome: This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions.
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method (2024.naacl-long)

Copied to clipboard

Challenge: Recent literature reveals that Large Language Models (LLMs) hallucinate intermittently, which impedes their reliability for further utilization.
Approach: They propose a self-detection method to detect which questions an LLM does not know by combining the two components to identify whether the model generates a non-factual response to the question.
Outcome: The proposed method can detect which questions an LLM does not know across factoid question-answering, arithmetic reasoning, and commonsense reasoning tasks.
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Working memory is a critical component of human intelligence and executive functioning . it is correlated with performance on various cognitive tasks, including fluid intelligence .
Approach: They apply working memory tasks to large language models to estimate working memory capacity . they find that LLMs exceed normative human scores, but not executive functioning benchmarks .
Outcome: The proposed models do not show higher performance on executive functioning tasks or problem solving benchmarks.
From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models excel at solving individual problems in isolation, but are they able to effectively collaborate over long-term interactions?
Approach: They propose to use a multi-session dataset to test LLMs' ability to track and execute simple coding instructions amid irrelevant information, simulating a realistic setting.
Outcome: The proposed model performs poorly when instructions are spread across sessions, suggesting that they are not able to integrate information over long interactions.
Substance Beats Style: Why Beginning Students Fail to Code with LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Existing work shows that beginners struggle to prompt LLMs to solve text-to-code tasks.
Approach: They propose to use a causal intervention experiment on technical vocabulary to test whether students lack the technical vocabulary needed to write good prompts and to analyze graphs that abstract how students edit prompts.
Outcome: The proposed model improves student-LLM communication by predicting student failures and predicting the information content of prompts.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations