AI, Take the Wheel: What Drives Delegation and Trust in Human–Computer Cooperative Question Answering? (2026.findings-acl)
Copied to clipboard
Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Zhu Irene Ying, Tianyi Zhou, Jordan Lee Boyd-Graber
| Challenge: | Human-AI collaboration is already happening, both in proactive delegation and deliberative adoption settings. |
| Approach: | They study delegating a task to AI without seeing its output and evaluating AI suggestions to decide whether to adopt them how AI output shapes final decisions. |
| Outcome: | The proposed game pairs 23 experts with 16 AI agents, capturing 387 delegation and 1440 adoption decisions. |
Similar Papers
Human-AI Collaboration: How AIs Augment Human Teammates (2025.acl-tutorials)
Copied to clipboard
| Challenge: | Despite the potential of general-purpose models, they are far from perfect, excelling at certain tasks while struggling with others. |
| Approach: | This tutorial will review recent developments related to human-AI teaming and collaboration. |
| Outcome: | This tutorial will review recent developments related to human-AI teaming and collaboration. |
A Diachronic Perspective on User Trust in AI under Uncertainty (2023.emnlp-main)
Copied to clipboard
| Challenge: | Modern NLP systems are rarely calibrated and are often confidently incorrect about their predictions, which violates users’ mental model and erodes their trust. |
| Approach: | They propose to use a mental model to bet on the correctness of an NLP system and to study how trust is rebuilt as a function of time after these events. |
| Outcome: | The proposed model shows that even a few highly inaccurate confidence estimation instances damage users’ trust in the system and performance, which does not easily recover over time. |
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies (2026.findings-acl)
Copied to clipboard
Myke C. Cohen, Mingqian Zheng, Neel Bhandari, Hsien-Te Kao, Xuhui Zhou, Daniel Nguyen, Laura Cassani, Maarten Sap, Svitlana Volkova
| Challenge: | In simulations, personality traits and AI attributes were comparatively influential, but with actual human subjects, AI attributes – particularly transparency – were much more impactful. |
| Approach: | They compare a purely simulated dataset and a parallel human subjects experiment to examine how human personality traits and AI design characteristics jointly shape interaction outcomes in imperfectly cooperative scenarios. |
| Outcome: | The results show that personality traits and AI attributes are comparatively influential in simulations, but with actual human subjects, they are much more impactful. |
Learning to Explain Selectively: A Case Study on Question Answering (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in machine learning (ML) have obstructed the use of NNs. |
| Approach: | They propose to learn to explain"selectively" for each decision that the user makes . they use a model to choose the best explanation from a set of candidates and update this model with feedback . |
| Outcome: | The proposed model improves human performance on a question-based task for experts and crowdworkers. |
On Evaluating Explanation Utility for Human-AI Decision Making in NLP (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a lack of evidence that explanations help people in situations they are introduced for is a problem in NLP . prior work on explainability has focused on overcoming technical challenges and used proxy evaluations. |
| Approach: | They propose to use existing metrics to evaluate the effectiveness of explanations in NLP . they argue that providing AI predictions does not cause decision makers to speed up work . |
| Outcome: | The proposed evaluations show that providing AI predictions does not cause decision makers to speed up their work without compromising performance. |
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond (2025.acl-long)
Copied to clipboard
| Challenge: | Using large language models, intelligent models have evolved into autonomous agents . this paradigm has yielded remarkable progress in numerous NLP tasks in recent years . |
| Approach: | They present a review of human-model cooperation, exploring its principles, formalizations, and open challenges. |
| Outcome: | The proposed model-model cooperation paradigm has been a key focus of recent research . it is a novel paradigm that can be applied to a variety of tasks . |
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models have led to claims of AI surpassing humans in QA tasks . authors: models are purportedly acing tests that many humans find challenging . |
| Approach: | They propose a framework that enables quantitative assessment and comparison of problem-solving abilities in QA agents. |
| Outcome: | The proposed framework uncovers distinctficiency patterns in knowledge domains and reasoning skills. |
Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution (2025.findings-emnlp)
Copied to clipboard
| Challenge: | despite growing interest in explainable NLP, it remains unclear how explanation strategies shape user behavior in tasks like authorship identification. |
| Approach: | They propose two explanation types to support their analysis of user behavior . they use example-based style rewrites and feature-based rationales to generate explanations . |
| Outcome: | The proposed explanations support appropriate reliance, whereas explanations increase AI overreliance, the study finds . |
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing systems that make predictions and ask questions are unable to have a mutual exchange of opinions. |
| Approach: | They propose to use a dataset and computational framework to allow systems to have beneficial discussions with humans, improving the accuracy by 25 points on a natural language inference task. |
| Outcome: | The proposed system improves accuracy by 25 points on a natural language inference task. |
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated (2025.findings-acl)
Copied to clipboard
| Challenge: | Prior research on AI mistrust focused primarily on AI's bias towards different human pop-ups. |
| Approach: | They examine how bias shapes the perception of AI versus human generated content . they found that raters favored content labeled "Human Generated" even when labels were deliberately swapped . |
| Outcome: | The findings highlight the limitations of human judgment in interacting with AI and offer a foundation for improving human-AI collaboration. |