Challenge: Human-AI collaboration is already happening, both in proactive delegation and deliberative adoption settings.
Approach: They study delegating a task to AI without seeing its output and evaluating AI suggestions to decide whether to adopt them how AI output shapes final decisions.
Outcome: The proposed game pairs 23 experts with 16 AI agents, capturing 387 delegation and 1440 adoption decisions.

Similar Papers

Human-AI Collaboration: How AIs Augment Human Teammates (2025.acl-tutorials)

Copied to clipboard

Challenge: Despite the potential of general-purpose models, they are far from perfect, excelling at certain tasks while struggling with others.
Approach: This tutorial will review recent developments related to human-AI teaming and collaboration.
Outcome: This tutorial will review recent developments related to human-AI teaming and collaboration.
A Diachronic Perspective on User Trust in AI under Uncertainty (2023.emnlp-main)

Copied to clipboard

Challenge: Modern NLP systems are rarely calibrated and are often confidently incorrect about their predictions, which violates users’ mental model and erodes their trust.
Approach: They propose to use a mental model to bet on the correctness of an NLP system and to study how trust is rebuilt as a function of time after these events.
Outcome: The proposed model shows that even a few highly inaccurate confidence estimation instances damage users’ trust in the system and performance, which does not easily recover over time.
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies (2026.findings-acl)

Copied to clipboard

Challenge: In simulations, personality traits and AI attributes were comparatively influential, but with actual human subjects, AI attributes – particularly transparency – were much more impactful.
Approach: They compare a purely simulated dataset and a parallel human subjects experiment to examine how human personality traits and AI design characteristics jointly shape interaction outcomes in imperfectly cooperative scenarios.
Outcome: The results show that personality traits and AI attributes are comparatively influential in simulations, but with actual human subjects, they are much more impactful.
Learning to Explain Selectively: A Case Study on Question Answering (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in machine learning (ML) have obstructed the use of NNs.
Approach: They propose to learn to explain"selectively" for each decision that the user makes . they use a model to choose the best explanation from a set of candidates and update this model with feedback .
Outcome: The proposed model improves human performance on a question-based task for experts and crowdworkers.
On Evaluating Explanation Utility for Human-AI Decision Making in NLP (2024.findings-emnlp)

Copied to clipboard

Challenge: a lack of evidence that explanations help people in situations they are introduced for is a problem in NLP . prior work on explainability has focused on overcoming technical challenges and used proxy evaluations.
Approach: They propose to use existing metrics to evaluate the effectiveness of explanations in NLP . they argue that providing AI predictions does not cause decision makers to speed up work .
Outcome: The proposed evaluations show that providing AI predictions does not cause decision makers to speed up their work without compromising performance.
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond (2025.acl-long)

Copied to clipboard

Challenge: Using large language models, intelligent models have evolved into autonomous agents . this paradigm has yielded remarkable progress in numerous NLP tasks in recent years .
Approach: They present a review of human-model cooperation, exploring its principles, formalizations, and open challenges.
Outcome: The proposed model-model cooperation paradigm has been a key focus of recent research . it is a novel paradigm that can be applied to a variety of tasks .
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have led to claims of AI surpassing humans in QA tasks . authors: models are purportedly acing tests that many humans find challenging .
Approach: They propose a framework that enables quantitative assessment and comparison of problem-solving abilities in QA agents.
Outcome: The proposed framework uncovers distinctficiency patterns in knowledge domains and reasoning skills.
Who’s the Author? How Explanations Impact User Reliance in AI-Assisted Authorship Attribution (2025.findings-emnlp)

Copied to clipboard

Challenge: despite growing interest in explainable NLP, it remains unclear how explanation strategies shape user behavior in tasks like authorship identification.
Approach: They propose two explanation types to support their analysis of user behavior . they use example-based style rewrites and feature-based rationales to generate explanations .
Outcome: The proposed explanations support appropriate reliance, whereas explanations increase AI overreliance, the study finds .
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach (2024.findings-eacl)

Copied to clipboard

Challenge: Existing systems that make predictions and ask questions are unable to have a mutual exchange of opinions.
Approach: They propose to use a dataset and computational framework to allow systems to have beneficial discussions with humans, improving the accuracy by 25 points on a natural language inference task.
Outcome: The proposed system improves accuracy by 25 points on a natural language inference task.
Human Bias in the Face of AI: Examining Human Judgment Against Text Labeled as AI Generated (2025.findings-acl)

Copied to clipboard

Challenge: Prior research on AI mistrust focused primarily on AI's bias towards different human pop-ups.
Approach: They examine how bias shapes the perception of AI versus human generated content . they found that raters favored content labeled "Human Generated" even when labels were deliberately swapped .
Outcome: The findings highlight the limitations of human judgment in interacting with AI and offer a foundation for improving human-AI collaboration.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations