Challenge: a growing number of people are seeking healthcare information from large language models via chatbots, yet the nature and inherent risks of these interactions remain unexplored.
Approach: They use a curated dataset of 11K real-world conversations composed of 25K user messages to analyze user interactions across 21 health specialties.
Outcome: The proposed dataset consists of 11K real-world conversations composed of 25K user messages.

Similar Papers

Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctness (2023.emnlp-main)

Copied to clipboard

Challenge: Using the TREC Misinformation dataset, we empirically evaluate ChatGPT to show not just its effectiveness but reveal that knowledge passed in the prompt can bias the model to the detriment of answer correctness.
Approach: They empirically evaluate ChatGPT to find out whether a prompt can bias the model to the detriment of answer correctness.
Outcome: The proposed model can be biased to the detriment of answer correctness by using retrieved-then-generate pipelines and how a user phrases their question as well as the question type.
Chatbot To Help Patients Understand Their Health (2025.findings-emnlp)

Copied to clipboard

Challenge: NoteAid-Chatbot is a conversational AI designed to help patients better understand their health .
Approach: They propose a new learning paradigm that leverages a multi-agent large language model and reinforcement learning framework without relying on costly human-generated training data.
Outcome: The proposed framework surpasses non-expert human training methods.
An “Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives” (2023.emnlp-main)

Copied to clipboard

Challenge: Mental health conversational agents (a.k.a. chatbots) are widely studied for their potential to offer accessible support to those experiencing mental health challenges.
Approach: They review 534 papers on building mental health-related conversational agents . they recommend a few recommendations to bridge the disciplinary divide .
Outcome: The systematic review reveals 136 key papers on building mental health-related conversational agents with diverse characteristics of modeling and experimental design techniques.
The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare Applications (2022.acl-long)

Copied to clipboard

Challenge: Task-oriented dialogue systems have been surveyed in the medical community from a non-technical perspective, but a systematic review from . a rigorous computational perspective has to date remained noticeably absent.
Approach: They analyze 4070 papers on task-oriented dialogue systems for healthcare applications and identify gaps in their analysis.
Outcome: The proposed system-level implementation details remain limited or underspecified, slowing the pace of innovation in this area.
Evaluating Large Language Models for Health-related Queries with Presuppositions (2024.findings-acl)

Copied to clipboard

Challenge: a large number of health-related queries require factually accurate answers . however, the lack of accurate answers may cause real-world harm .
Approach: They evaluate the factual accuracy and consistency of large language models using a dataset consisting of health-related queries with varying degrees of presuppositions.
Outcome: The proposed model responses agree with 23-32% of existing false claims and 49-55% with novel fabricated claims.
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models lack adequate evaluations and prompting strategies for explainability.
Approach: They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks.
Outcome: The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods.
Risk-graded Safety for Handling Medical Queries in Conversational AI (2022.aacl-short)

Copied to clipboard

Challenge: Conversational AI systems can engage in unsafe behaviour when handling medical queries that could lead to death.
Approach: They label medical queries with crowdsourced and expert annotations to identify the seriousness of the prompts and recognise the risk types posed by the responses.
Outcome: The results suggest that these tasks can be automated, but caution should be exercised, as errors can potentially be very serious.
NoteChat: A Dataset of Synthetic Patient-Physician Conversations Conditioned on Clinical Notes (2024.findings-acl)

Copied to clipboard

Challenge: NoteChat is a cooperative multi-agent framework for generating patient-physician dialogues . evaluator finds it outperforms state-of-the-art models for generating clinical notes . clinical documentation is largely done by physicians at both steps .
Approach: They propose a cooperative multi-agent framework leveraging Large Language Models to generate patient-physician dialogues.
Outcome: The proposed framework outperforms state-of-the-art models for generating clinical notes . it can engage patients directly and help clinical documentation, a leading cause of physician burnout .
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking (2026.findings-acl)

Copied to clipboard

Challenge: Creating spoken dialogue datasets is methodologically challenging due to the personally identifiable nature of speech signals.
Approach: They propose a large-scale, multilingual, and multi-parallel dataset for developing and evaluating retrieval-augmented generation-based spoken dialogue systems.
Outcome: The proposed dataset includes 6,000 information-seeking dialogues and 163 hours of user speech recorded from native speakers of four official WHO languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations