Challenge: Prior studies have shown that sufficient collaboration is the key factor that determines the outcome of an operation.
Approach: They propose to model the communication between team members during an operation using audio data and physiology signals from two camera angles.
Outcome: The proposed model is based on existing frameworks and invites future effort on developing methods that can deal with real-world clinical data.

Similar Papers

Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
NoteChat: A Dataset of Synthetic Patient-Physician Conversations Conditioned on Clinical Notes (2024.findings-acl)

Copied to clipboard

Challenge: NoteChat is a cooperative multi-agent framework for generating patient-physician dialogues . evaluator finds it outperforms state-of-the-art models for generating clinical notes . clinical documentation is largely done by physicians at both steps .
Approach: They propose a cooperative multi-agent framework leveraging Large Language Models to generate patient-physician dialogues.
Outcome: The proposed framework outperforms state-of-the-art models for generating clinical notes . it can engage patients directly and help clinical documentation, a leading cause of physician burnout .
HealthAlignSumm : Utilizing Alignment for Multimodal Summarization of Code-Mixed Healthcare Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: Collaboration between doctors and AI scientists is leading to personalized models to stream-line healthcare tasks and improve productivity.
Approach: They propose to use alignment techniques to combine a doctor-patient dialogue with a visual component of the BART model.
Outcome: The proposed model in-tegrates visual components with the BART ar-chitecture.
Multi-party Multimodal Conversations Between Patients, Their Companions, and a Social Robot in a Hospital Memory Clinic (2024.eacl-demo)

Copied to clipboard

Challenge: a new spoken dialogue system is being developed for hospitals and hospitals to enable multi-party interactions . a social robot can be used to have multi-part conversations with patients and their companions .
Approach: They describe a spoken dialogue system that allows patients to have multi-party conversations with their companions . they use speech and video input to generate both speech and gestures - arm, head, and eye movements .
Outcome: The proposed system generates human-like clarification requests when the patient pauses mid-utterance, answers in-domain questions, and responds appropriately to out-of-domain requests.
The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare Applications (2022.acl-long)

Copied to clipboard

Challenge: Task-oriented dialogue systems have been surveyed in the medical community from a non-technical perspective, but a systematic review from . a rigorous computational perspective has to date remained noticeably absent.
Approach: They analyze 4070 papers on task-oriented dialogue systems for healthcare applications and identify gaps in their analysis.
Outcome: The proposed system-level implementation details remain limited or underspecified, slowing the pace of innovation in this area.
From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models excel at solving individual problems in isolation, but are they able to effectively collaborate over long-term interactions?
Approach: They propose to use a multi-session dataset to test LLMs' ability to track and execute simple coding instructions amid irrelevant information, simulating a realistic setting.
Outcome: The proposed model performs poorly when instructions are spread across sessions, suggesting that they are not able to integrate information over long interactions.
Common Ground Tracking in Multimodal Dialogue (2024.lrec-main)

Copied to clipboard

Challenge: In dialogue modeling, there is considerable attention on “dialogue state tracking” (DST) but “common ground tracking” identifies the shared belief space held by all participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true.
Approach: They propose a method for automatically identifying the current set of shared beliefs and ”questions under discussion” of a group with a shared goal.
Outcome: The proposed method predicts moves toward building common ground relative to ground truth in a multimodal interaction with an AI.
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
MONAH: Multi-Modal Narratives for Humans to analyze conversations (2021.eacl-main)

Copied to clipboard

Challenge: In conversational analyses, humans manually weave multimodal information into the transcripts, which is significantly time-consuming.
Approach: They propose a system that automatically expands the verbatim transcripts of video-recorded conversations using multimodal data streams.
Outcome: The proposed system improves detecting rapport-building by expanding the range of multimodal annotations.
Interactive Evaluation for Medical LLMs via Task-oriented Dialogue System (2025.coling-main)

Copied to clipboard

Challenge: In typical medical scenarios, doctors often ask a set of questions to gain a comprehensive understanding of patients’ conditions.
Approach: They propose to use multi-turn medical dialogue evaluation to evaluate proactive communication and diagnostic capabilities of medical Large Language Models (LLMs) .
Outcome: The proposed model outperforms existing models on multi-turn question-answering datasets and is therefore cost-effective.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations