Challenge: Existing studies on large language models (LLMs) fail to detect character knowledge errors, leading to low-quality automatic corpus construction.
Approach: They propose to use a large language model to detect known knowledge errors and an agent-based reasoning method to improve error detection.
Outcome: The proposed method improves the ability of LLMs to detect errors in known knowledge errors and unknown knowledge errors while playing roles.

Similar Papers

Do LLMs Catch Their Own Mistakes? A Comprehensive Benchmark for Reflective Tool Use LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks primarily evaluate planning and execution success, overlooking the self-reflective dimension of tool use.
Approach: They propose a benchmark to assess LLMs’ self-reflective reasoning in tool-augmented multi-turn dialogues.
Outcome: The proposed benchmark covers 10 domains with 88 distinct APIs and 968 annotated dialogues, systematically injecting diverse error types arising from both user and assistant behavior.
LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting missing foreign keys are limited in capturing semantic dependencies across schemas.
Approach: They propose a framework that integrates four agents to detect missing foreign keys . they propose combinatorial search space explosion, ambiguous inference and global inconsistency .
Outcome: The proposed framework achieves F1-scores above 93% on large-scale MusicBrainz database . it reduces candidate search space by two to three orders of magnitude without losing true FKs .
Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing detection methods fail to account for **self-consistent error** . study identifies self-consistency errors and evaluates them .
Approach: They propose a method that fuses hidden state evidence from an external verifier LLM to detect self-consistent errors.
Outcome: The proposed method significantly enhances performance on self-consistent errors across three LLM families.
LLMs cannot find reasoning errors, but can correct them given the error location (2024.findings-acl)

Copied to clipboard

Challenge: Recent attempts to self-correct logical or reasoning errors often cause correct answers to become incorrect, resulting in poor performance overall.
Approach: They propose to use a backtracking setup to test the correction abilities of LLMs on their mistake-finding ability to find logical mistakes.
Outcome: The proposed model improves on 5 reasoning tasks, showing that it can correct logical mistakes without ground truth labels or training data.
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method (2024.naacl-long)

Copied to clipboard

Challenge: Recent literature reveals that Large Language Models (LLMs) hallucinate intermittently, which impedes their reliability for further utilization.
Approach: They propose a self-detection method to detect which questions an LLM does not know by combining the two components to identify whether the model generates a non-factual response to the question.
Outcome: The proposed method can detect which questions an LLM does not know across factoid question-answering, arithmetic reasoning, and commonsense reasoning tasks.
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: CriticBench is a benchmark designed to assess LLMs’ abilities to critique and refine their reasoning across a variety of tasks.
Approach: They propose a benchmark to assess LLMs' ability to critique and correct reasoning across a variety of tasks.
Outcome: The proposed benchmark examines the performance of 17 large language models in generation, critique, and correction reasoning.
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios (2025.emnlp-main)

Copied to clipboard

Challenge: a number of tools are used to perform complex tasks, but the tool utilization process can cause errors.
Approach: They propose a critique evaluation benchmark for tool learning that analyzes function-calling errors on tool evaluation benchmarks.
Outcome: The proposed critique evaluation benchmark holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios.
Towards Self-Improving Error Diagnosis in Multi-Agent Systems (2026.findings-acl)

Copied to clipboard

Challenge: Existing diagnostic approaches rely on expensive expert annotations and ”LLM-as-a-judge” paradigms.
Approach: They propose a framework for semantic failure attribution that identifies responsible agents and the originating error step.
Outcome: The proposed framework outperforms baselines in step-level localization and validation.
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations.
Approach: They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions .
Outcome: Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations