TestAgent: An Adaptive and Intelligent Expert for Human Assessment (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing adaptive testing methods face several challenges due to mechanized nature of most algorithms and noisy response data. |
| Approach: | They propose to use large language models to enhance adaptive testing through interactive engagement to capture test-takers’ responses and anomalies. |
| Outcome: | The proposed agent achieves more accurate results with 20% fewer questions than state-of-the-art baselines and testers preferred it in speed, smoothness, and other dimensions. |
Similar Papers
Towards Explainable Computerized Adaptive Testing with Large Language Model (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods focus on minimizing the number of questions required to assess ability, lacking clear and reliable explanations for the question selection process. |
| Approach: | They propose to use large language models to enhance computer adaptive testing (CAT) by providing human-like interpretability and explanations. |
| Outcome: | The proposed agent-based CAT performs comparably or superior to traditional CAT methods in accuracy and significantly improves student trust and satisfaction. |
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation (2026.findings-acl)
Copied to clipboard
Zhichao Shi, Xuhui Jiang, Chengjin Xu, Cangli Yao, Shengjie Ma, Yinghan Shen, Zixuan Li, Jian Guo, Yuanzhuo Wang
| Challenge: | Current evaluation methods for large language models rely on static benchmarks . limited knowledge coverage and fixed difficulties hinder the targeted optimizations resulting in superficial evaluations of LLMs - a problem that has been addressed by JudgeAgent . |
| Approach: | They propose a knowledge-driven and dynamic evaluation framework for large language models . judgeAgent leverages LLM agents equipped with context graphs to traverse knowledge structures . |
| Outcome: | The proposed framework can achieve comprehensive evaluations and facilitate effective model iterations. |
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents (2026.eacl-long)
Copied to clipboard
| Challenge: | Robust, developer-friendly evaluation remains a bottleneck. |
| Approach: | They propose a meta-agent that combines static code analysis, developer interrogation, literature mining, and persona-driven adversarial test generation whose difficulty adapts via judge feedback. |
| Outcome: | The agent-testing agent (ATA) surfaces more diverse and severe failures than expert annotators while matching severity, and finishes in 20–30 minutes versus ten-annotator rounds that took days. |
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub (2025.acl-long)
Copied to clipboard
Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun
| Challenge: | Existing approaches lack flexibility to address diverse and ever-evolving user queries in open domains. |
| Approach: | They propose to evaluate LLMs on open-domain knowledge that requires tools to solve diverse and ever-evolving user queries. |
| Outcome: | The proposed system outperforms baselines in the open domain task-solving benchmark. |
PersonaAgent: Bridging Memory and Action for Personalized LLM Agents (2026.findings-acl)
Copied to clipboard
Weizhi Zhang, Xinyang Zhang, Chenwei Zhang, Liangwei Yang, Jingbo Shang, Zhepei Wei, Henry Peng Zou, Zijie Huang, Zhengyang Wang, Yifan Gao, Xiaoman Pan, Lian Xiong, Jingguo Liu, Philip S. Yu, Xian Li
| Challenge: | Existing Large Language Model (LLM) enabled agents lack flexibility to respond to users’ varying needs and preferences. |
| Approach: | They propose a test-time user-preference alignment strategy that optimizes the persona prompt, ensuring real-time preference alignment through textual loss feedback between simulated and ground-truth responses. |
| Outcome: | The proposed framework outperforms baseline methods in real-time and in real applications. |
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models (2026.findings-acl)
Copied to clipboard
Bang Zhang, Ruotian Ma, Qingxuan Jiang, Peisong Wang, Jiaqi Chen, Zheng Xie, Xingyu Chen, Yue Wang, Fanghua Ye, Jian Li, Yifan Yang, Zhaopeng Tu, Xiaolong Li
| Challenge: | Large language models (LLMs) have evolved from statistical sequence predictors to sophisticated autonomous agents capable of reasoning, planning, and sustaining multi-turn conversa-tions. |
| Approach: | They propose a system that instantiates a "Sentient Agent" that simulates human-like emotional changes and inner thoughts to provide a more realistic evaluation of the model in multi-turn conversations. |
| Outcome: | The proposed framework measures the agent's higher-order social cognition in multi-turn conversations. |
Agentic AI for Human Resources: LLM-Driven Candidate Assessment (2026.eacl-demo)
Copied to clipboard
Kamer Ali Yuksel, Abdul Basit Anees, Ashraf Hatim Elneima, Sanjika Hewavitharana, Mohamed Al-Badrashiny, Hassan Sawaf
| Challenge: | Current systems rely on keyword matching and shallow keyword-based screening, leading to missed opportunities and inconsistent evaluations. |
| Approach: | They propose a framework that uses Large Language Models to automate candidate assessment in recruitment. |
| Outcome: | The proposed framework outputs detailed assessment reports, candidate comparisons, and ranked recommendations that are transparent, auditable, and suitable for real-world hiring workflows. |
Machine Learning–Driven Language Assessment (2020.tacl-1)
Copied to clipboard
| Challenge: | Language proficiency tests are cumbersome to create and maintain, and items may be copied and leaked or simply used too often. |
| Approach: | They propose a method that uses machine learning and natural language processing to induce proficiency scales and linguistic models to estimate item difficulty directly for computer-adaptive testing. |
| Outcome: | The proposed method produces scores that are reliable and reliable while generating item banks large enough to satisfy security requirements. |
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks primarily assess static knowledge, while intelligence also entails the ability to rapidly learn from experience. |
| Approach: | They propose to use semantic games to evaluate test-time learning . they recruit eight human participants to complete the same task . |
| Outcome: | The proposed framework compares model performance under limited and cumulative experience settings and contains four forms of experience representation. |
PersonaGym: Evaluating Persona Agents and LLMs (2025.findings-emnlp)
Copied to clipboard
Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik R Narasimhan, Vishvak Murahari
| Challenge: | Persona agents are LLM agents conditioned to act according to an assigned persona . evaluating how faithfully these agents adhere to their personas remains a challenge . |
| Approach: | a new study evaluates persona agents' ability to act according to an assigned persona . a persona agent's person score is a human-aligned automatic metric that can be used to evaluate a model . |
| Outcome: | a new evaluation framework and a human-aligned automatic metric show that persona agents can perform better. |