Challenge: Existing evaluation benchmarks, such as MMLU, C-Eval, and GSM8K, evaluate models by posing a variety of problems, including problems about mathematics, science, law, and general knowledge.
Approach: They propose a benchmark which assesses the model’s lateral thinking within an interactive framework.
Outcome: The proposed evaluation benchmark assesses the model’s lateral thinking within an interactive framework.

Similar Papers

FineReason: Evaluating and Improving LLMs’ Deliberate Reasoning through Reflective Puzzle Solving (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) highlight an important shift from the “System 1” way of quick reactions to the “system 2” style of reflection-and-correction problem solving.
Approach: They propose a logic-puzzle benchmark for systematic evaluation of large language models' reasoning capabilities that decomposes each puzzle into atomic steps.
Outcome: The proposed model improves on state checking and state transition tasks and demonstrates gains in reasoning by up to 5.1%.
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation.
Approach: They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching.
Outcome: The proposed model performs poorly when faced with reordered constraints or irrelevant information.
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have achieved impressive progress on many benchmarks, yet they still have fundamental weaknesses.
Approach: They introduce Concept, a word-guessing board game, as a benchmark for probing abductive reasoning.
Outcome: The proposed game is easily solved by humans, but is still very challenging for state-of-the-art LLMs (no model exceeds 40% success rate).
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.
LLM-Evolve: Evaluation for LLM’s Evolving Capability on Benchmarks (2024.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models evaluate LLMs on i.i.d. tasks, overlooking their ability to learn iteratively from past experiences.
Approach: They propose a framework which extends established benchmarks to sequential problem-solving settings and provides feedback after each round to build a demonstration memory that the models can query in future tasks.
Outcome: The proposed framework can improve performance of LLMs by learning from past interactions and improve models' performance over time.
Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive Scenarios (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models fail to recognize fallacious reasoning in real-world interactions despite strong performance on static fallacy detection tasks.
Approach: They propose a Chinese benchmark to assess fallacy awareness without explicit cues . they propose 'fate' evaluation framework that assesses fallacy without explicit .
Outcome: The proposed framework assesses fallacy awareness without explicit cues, combining natural dialogue responses and reasoning-based decisions.
SATBench: Benchmarking LLMs’ Logical Reasoning via Automated Puzzle Generation from SAT Formulas (2025.emnlp-main)

Copied to clipboard

Challenge: SATBench is a benchmark for evaluating the logical reasoning capabilities of large language models (LLMs) through logical puzzles derived from Boolean satisfiability (SAT) problems.
Approach: They propose a benchmark to evaluate logical reasoning capabilities of large language models (LLMs) using logical puzzles derived from Boolean satisfiability problems.
Outcome: The proposed model achieves 65.0% accuracy on hard UNSAT problems, close to the baseline of 50%.
From Informal to Formal – Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs (2025.acl-long)

Copied to clipboard

Challenge: Recent studies in formal mathematical reasoning have shown an unstoppable growth trend.
Approach: They constructed 18k high-quality instruction-response pairs across five mainstream formal specification languages and evaluated them against ten open-sourced LLMs.
Outcome: The proposed model compared instruction-response pairs across five formal specification languages and found that the LLMs were good at writing proof segments when given either the code, or the detailed description of proof steps.
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing logical reasoning evaluation benchmarks focus on simplistic single-step or multi-step reasoning with limited set of inference rules.
Approach: They propose to use a multi-step logical reasoning evaluation dataset to measure their ability for human-like multi- step logical thinking.
Outcome: The proposed dataset covers three logic types including propositional, first-order, and non-monotonic logic with various inference rules and depths.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations