SMILEE: Symmetric Multi-modal Interactions with Language-gesture Enabled (AI) Embodiment (N18-5)
Copied to clipboard
| Challenge: | SMILEE is a conversational agent system that interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures. |
| Approach: | They propose to use a computer-generated avatar to embody a human-machine conversational agent system that interprets verbal utterances and non-verbal behaviors to facilitate natural symmetric multi-modal interactions. |
| Outcome: | The proposed system interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures, and communicates with natural language and gestures through its embodiment as an avatar. |
Similar Papers
Cue-bot: A Conversational Agent for Assistive Technology (2022.acl-demo)
Copied to clipboard
Shachi H Kumar, Hsuan Su, Ramesh Manuvinakurike, Maximilian C. Pinaroc, Sai Prasad, Saurav Sahay, Lama Nachman
| Challenge: | Large-scale pre-training has achieved significant performance gains across many tasks within NLP, including intent prediction and dialogue state tracking. |
| Approach: | They propose to use eye-tracking, mouse controls and an intelligent agent Cue-bot to represent the user in a conversation. |
| Outcome: | The proposed system can be used by people with different levels of disabilities to interact with the world, supported by eye-tracking, mouse controls and an intelligent agent Cue-bot. |
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size. |
| Approach: | They combine open-domain dialogue agents with vision models to investigate human preferences and humanness. |
| Outcome: | The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot. |
AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Recent advances in AI focus on multi-agent systems (MAS) that can be integrated with Large Language Models (LLMs) but current systems still face challenges of inter-agency communication, coordination, and interaction with heterogeneous tools and resources. |
| Approach: | They propose a modular multi-protocol MAS framework with self-implemented A2A and MCP . the framework supports natural language interaction without prior technical expertise . |
| Outcome: | The proposed framework supports natural language interaction without prior technical expertise and responds to multimodal queries for tasks including information retrieval, question answering, and image analysis. |
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
| Approach: | This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
| Outcome: | This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
An Evaluation Framework for Multimodal Interaction (L18-1)
Copied to clipboard
| Challenge: | a framework for evaluating multimodal interactions is presented . it leverages the semantics of language and gesture to assess mutual understanding . consistent evaluation is required to test areas where the system needs improvement . |
| Approach: | They propose a framework for evaluating interactions between human and virtual agent . they use VoxML as a platform to model interactions using natural language and gesture . |
| Outcome: | The proposed framework assesses the level of mutual understanding and ease of communication between human and computer agents in a blocks world scenario. |
tagE: Enabling an Embodied Agent to Understand Human Instructions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing systems for natural language understanding (NLU) are limited due to the inherent ambiguity and incompleteness inherent in natural language. |
| Approach: | They propose a system to extract tasks from natural language instructions and map them to robots' established collection of skills. |
| Outcome: | The proposed system outperforms baseline models in the training and evaluation of a dataset featuring complex instructions. |
Spoken Conversational Agents with Large Language Models (2025.emnlp-tutorials)
Copied to clipboard
| Challenge: | This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training . |
| Approach: | This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems . |
| Outcome: | This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents . |
Achieving Common Ground in Multi-modal Dialogue (2020.acl-tutorials)
Copied to clipboard
| Challenge: | tutorial focuses on three main topic areas: grounding in human-human communication, dialogue systems and multi-modal interactive systems. |
| Approach: | This tutorial examines the use of computational dialogue research to design grounding modules and behaviors in cutting-edge systems. |
| Outcome: | This tutorial examines the results of recent research on grounding in human-human communication . it shows how these results lead to rich and challenging opportunities for doing grounding more flexible and powerful ways . |
AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Large language models (LLMs) are being used for planning in orchestrated multi-agent systems . existing LLMs fall short of human expectations and lack effective mechanisms for users to inspect, understand, and control their behaviors. |
| Approach: | They propose a system supporting human-in-the-loop planning through conversational and graph-based interfaces. |
| Outcome: | AIPOM enables users to transparently inspect, refine, and collaboratively guide LLM-generated plans, significantly enhancing user control and trust in multi-agent workflows. |
ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents (2020.acl-demos)
Copied to clipboard
Chia-Yu Li, Daniel Ortega, Dirk Väth, Florian Lux, Lindsey Vanderlyn, Maximilian Schmidt, Michael Neumann, Moritz Völkel, Pavel Denisov, Sabrina Jenne, Zorica Kacarevic, Ngoc Thang Vu
| Challenge: | Existing toolkits for developing dialog systems are limited to core components and do not support multi-modal processing and social signals. |
| Approach: | They propose to use ADVISER to develop multi-modal dialog agents using multi-text and social signals. |
| Outcome: | The proposed toolkit is flexible, easy to use, and easy to extend for linguists and cognitive scientists, thereby providing a flexible platform for collaborative research. |