Challenge: SMILEE is a conversational agent system that interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures.
Approach: They propose to use a computer-generated avatar to embody a human-machine conversational agent system that interprets verbal utterances and non-verbal behaviors to facilitate natural symmetric multi-modal interactions.
Outcome: The proposed system interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures, and communicates with natural language and gestures through its embodiment as an avatar.

Similar Papers

Cue-bot: A Conversational Agent for Assistive Technology (2022.acl-demo)

Copied to clipboard

Challenge: Large-scale pre-training has achieved significant performance gains across many tasks within NLP, including intent prediction and dialogue state tracking.
Approach: They propose to use eye-tracking, mouse controls and an intelligent agent Cue-bot to represent the user in a conversation.
Outcome: The proposed system can be used by people with different levels of disabilities to interact with the world, supported by eye-tracking, mouse controls and an intelligent agent Cue-bot.
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size.
Approach: They combine open-domain dialogue agents with vision models to investigate human preferences and humanness.
Outcome: The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot.
AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis (2025.emnlp-demos)

Copied to clipboard

Challenge: Recent advances in AI focus on multi-agent systems (MAS) that can be integrated with Large Language Models (LLMs) but current systems still face challenges of inter-agency communication, coordination, and interaction with heterogeneous tools and resources.
Approach: They propose a modular multi-protocol MAS framework with self-implemented A2A and MCP . the framework supports natural language interaction without prior technical expertise .
Outcome: The proposed framework supports natural language interaction without prior technical expertise and responds to multimodal queries for tasks including information retrieval, question answering, and image analysis.
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Approach: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Outcome: This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models.
An Evaluation Framework for Multimodal Interaction (L18-1)

Copied to clipboard

Challenge: a framework for evaluating multimodal interactions is presented . it leverages the semantics of language and gesture to assess mutual understanding . consistent evaluation is required to test areas where the system needs improvement .
Approach: They propose a framework for evaluating interactions between human and virtual agent . they use VoxML as a platform to model interactions using natural language and gesture .
Outcome: The proposed framework assesses the level of mutual understanding and ease of communication between human and computer agents in a blocks world scenario.
tagE: Enabling an Embodied Agent to Understand Human Instructions (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for natural language understanding (NLU) are limited due to the inherent ambiguity and incompleteness inherent in natural language.
Approach: They propose a system to extract tasks from natural language instructions and map them to robots' established collection of skills.
Outcome: The proposed system outperforms baseline models in the training and evaluation of a dataset featuring complex instructions.
Spoken Conversational Agents with Large Language Models (2025.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training .
Approach: This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems .
Outcome: This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents .
Achieving Common Ground in Multi-modal Dialogue (2020.acl-tutorials)

Copied to clipboard

Challenge: tutorial focuses on three main topic areas: grounding in human-human communication, dialogue systems and multi-modal interactive systems.
Approach: This tutorial examines the use of computational dialogue research to design grounding modules and behaviors in cutting-edge systems.
Outcome: This tutorial examines the results of recent research on grounding in human-human communication . it shows how these results lead to rich and challenging opportunities for doing grounding more flexible and powerful ways .
AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems (2025.emnlp-demos)

Copied to clipboard

Challenge: Large language models (LLMs) are being used for planning in orchestrated multi-agent systems . existing LLMs fall short of human expectations and lack effective mechanisms for users to inspect, understand, and control their behaviors.
Approach: They propose a system supporting human-in-the-loop planning through conversational and graph-based interfaces.
Outcome: AIPOM enables users to transparently inspect, refine, and collaboratively guide LLM-generated plans, significantly enhancing user control and trust in multi-agent workflows.
ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents (2020.acl-demos)

Copied to clipboard

Challenge: Existing toolkits for developing dialog systems are limited to core components and do not support multi-modal processing and social signals.
Approach: They propose to use ADVISER to develop multi-modal dialog agents using multi-text and social signals.
Outcome: The proposed toolkit is flexible, easy to use, and easy to extend for linguists and cognitive scientists, thereby providing a flexible platform for collaborative research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations