Challenge: Large language model agents exhibit action boundary blindness, granularity confusion, scope creep and boundary ambiguity . Explicit boundary prompting improves ABS by 0.08–0.13 across all models .
Approach: They propose four automatic metrics that require no human annotation to detect boundary blindness . they propose to use a multi-label attribution framework to validate the models .
Outcome: Experiments with seven large language model agents show that the best model achieves only 0.424 ABS . Explicit Boundary Prompting improves ABS by 0.08–0.13 across all models .

Similar Papers

Beyond Blind Following: Evaluating Robustness of LLM Agents under Imperfect Guidance (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown strong capabilities as task-solving agents across interactive domains, but in complex environments, auxiliary guidance may be imperfect.
Approach: They propose a benchmark to measure the robustness of large language models under imperfect guidance.
Outcome: The proposed benchmark compared LLM agents in navigation, cooking, and gaming in a variety of environments with auxiliary guidance and noisy or underspecified instructions extracted from demonstrations.
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have remarkable performance in a variety of tasks due to factual knowledge accumulated during pre-training.
Approach: They propose an evaluation protocol that disentangles knowing from applying knowledge and test state-of-the-art LLMs on 49 ambiguous entities.
Outcome: The proposed evaluation protocol disentangles knowing from applying knowledge and tests state-of-the-art LLMs on 49 ambiguous entities.
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments (2025.findings-emnlp)

Copied to clipboard

Challenge: Experiments with 4 different LLMs across 5 embodied environments show significant efficiency improvements, with only minor drops in agent performance.
Approach: They propose an intrinsic method that injects exit instructions during generation and an extransic system that verifies task completion to determine when to halt an agent’s trial.
Outcome: The proposed method injects exit instructions during generation and an exit method verifies task completion to determine when to halt an agent’s trial.
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception (2026.findings-acl)

Copied to clipboard

Challenge: Large language model agents assume a stationary context, failing to account for real-world time elapsed between messages.
Approach: They construct a dataset of multi-turn user–agent message trajectories across 76 scenarios . they collect human preferences between "calling a tool" and "directly answering" they also examine whether existing models lack human temporal perception .
Outcome: The results show that existing models display poor alignment with human temporal perception . the findings provide insights to foster the development of more time-aware and human-aligned agents.
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit impressive performance across diverse tasks but struggle to accurately gauge their knowledge boundaries.
Approach: They propose Consistency-based Confidence Calibration (C3) which assesses confidence consistency through question reformulation to improve LLMs’ ability to recognize their knowledge gaps.
Outcome: The proposed method improves the unknown perception rate by 5.6% on NQ and 4.9% on HotpotQA.
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability.
Approach: They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria.
Outcome: The proposed system is based on 11 common aspects with different evaluation criteria.
Can a Large Language Model Keep My Secrets? A Study on LLM-Controlled Agents (2025.acl-srw)

Copied to clipboard

Challenge: Using large language models, agents can assist with natural language tasks when given access to confidential data.
Approach: They created a synthetic dataset consisting of confidentiality-aware planning and deduction tasks in organizational access control.
Outcome: The proposed model can perform tasks similar to humans when given access to confidential data.
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have been developed to deal with real-world crimes, but it remains unclear whether they internalize authentic knowledge or are forced to simulate toxic language patterns.
Approach: They construct knowledge-intensive Q&A to investigate misuse threats of Large Language Models in terms of dangerous knowledge possession, harmful task planning utility, and harmfulness judgment robustness.
Outcome: The findings raise concerns that jailbreak success is often attributable to a hallucination loop between jailbroken LLM and judger LLM .
Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) often refuse to answer legitimate queries, causing models to treat many reasonable prompts as potentially risky.
Approach: They propose a framework that automatically generates and selects overrefusal prompts near the safety boundary.
Outcome: The proposed framework identifies and curates boundary-aligned prompts, enabling more effective and targeted mitigation of overrefusal.
Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations (2026.acl-long)

Copied to clipboard

Challenge: Existing static strategies for mitigating hallucinations do not explicitly model the information gain from interacting with the external environment.
Approach: They propose a calibration-driven interactive learning strategy that selects clarification queries by optimizing calibration error.
Outcome: The proposed method provides theoretical guarantees and empirical gains for reliability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations