Process Evaluation for Agentic Systems (2026.findings-eacl)

Copied to clipboard

Challenge: Recent adoption of LLM-based assistants has led to premature assumptions about their reliability and general capability.
Approach: They propose to assess the feasibility of automatic process evaluation for critical applications such as medicine, finance, law and infrastructure.
Outcome: The proposed evaluations are based on a small-scale study to assess the feasibility of automated process evaluation, present a compliance score, analyse use cases of bad and good behaviours, and offer recommendations for more holistic evaluation.

Similar Papers

A Survey on Evaluation of LLM-based Agents (2026.findings-acl)

Copied to clipboard

Challenge: This paper provides the first comprehensive survey of evaluation methods for LLM-based agents . LLMs are static, having fixed knowledge, and confined to text-to-text interaction.
Approach: They analyze the evaluation of LLM-based agents across five perspectives . they identify current trends and key gaps in evaluation methods .
Outcome: The proposed evaluation frameworks and tools are based on five perspectives . the results highlight current trends and identify gaps in future research .
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents (2026.acl-demo)

Copied to clipboard

Challenge: Agentic systems are becoming more capable of defining strategies, taking actions, and solving complex, multi-step tasks.
Approach: They propose an automatic, dynamic, and easy-to-use evaluation framework that provides textual insights into agent behavior on three levels of granularity: system, trace, and node.
Outcome: The proposed framework produces high-quality, data-driven, insightful feedback on system, trace, and node.
An Evaluation Mechanism of LLM-based Agents on Manipulating APIs (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable capabilities across a variety of tasks, such as language, mathematics, coding, and etc.
Approach: They propose to decompose tool use capability into seven aspects and form a thorough evaluation schema for generic agents.
Outcome: The proposed agent acts like a super-APP and can manipulate API-based tools.
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks evaluate agents in simplified, idealized settings, relying on pre-packaged tool interfaces, overlooking critical steps, and assume inputs are clean and fully specified.
Approach: They propose a framework that evaluates language agents in simplified, idealized settings . they show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 .
Outcome: Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 .
Agentic AI for Human Resources: LLM-Driven Candidate Assessment (2026.eacl-demo)

Copied to clipboard

Challenge: Current systems rely on keyword matching and shallow keyword-based screening, leading to missed opportunities and inconsistent evaluations.
Approach: They propose a framework that uses Large Language Models to automate candidate assessment in recruitment.
Outcome: The proposed framework outputs detailed assessment reports, candidate comparisons, and ranked recommendations that are transparent, auditable, and suitable for real-world hiring workflows.
PersonaGym: Evaluating Persona Agents and LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Persona agents are LLM agents conditioned to act according to an assigned persona . evaluating how faithfully these agents adhere to their personas remains a challenge .
Approach: a new study evaluates persona agents' ability to act according to an assigned persona . a persona agent's person score is a human-aligned automatic metric that can be used to evaluate a model .
Outcome: a new evaluation framework and a human-aligned automatic metric show that persona agents can perform better.
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents.
Approach: They propose to integrate human-provided information, feedback, or control into the agent system to enhance system performance, reliability, and safety.
Outcome: The proposed systems improve system performance, reliability, and safety by integrating human-provided information, feedback, or control into the agent system.
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments (2025.emnlp-main)

Copied to clipboard

Challenge: Enterprise systems are crucial for enhancing productivity and strategic growth, but data is fragmented across multiple sources and access controls are complex.
Approach: They propose a benchmark that simulates enterprise settings with 500 diverse tasks . they show that even the most capable models achieve only 41.8% task completion .
Outcome: The proposed benchmark shows that even the most capable models achieve only 41.8% task completion.
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for LLM-based mobile agents are insufficient to evaluate their capabilities.
Approach: They propose a benchmark to evaluate LLM-based mobile agents' planning capabilities . they expand UI operations by incorporating 103 APIs to accelerate task completion .
Outcome: The proposed benchmarks are based on 103 collected APIs and real user queries . the data is categorized into three distinct groups: SAST, SAMT, and MAMT .
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction (2026.findings-acl)

Copied to clipboard

Challenge: Current benchmarks evaluate task accuracy but overlook how agents interact . Preference-aware agents show 7.6% average UX improvement and 18.5% gain in preference alignment.
Approach: They propose a configurable environment that evaluates both what agents accomplish and how they interact.
Outcome: The proposed model improves performance and improves user experience by 7.6% and 18.5% respectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations