| Challenge: | a framework for evaluating multimodal interactions is presented . it leverages the semantics of language and gesture to assess mutual understanding . consistent evaluation is required to test areas where the system needs improvement . |
| Approach: | They propose a framework for evaluating interactions between human and virtual agent . they use VoxML as a platform to model interactions using natural language and gesture . |
| Outcome: | The proposed framework assesses the level of mutual understanding and ease of communication between human and computer agents in a blocks world scenario. |
Similar Papers
SMILEE: Symmetric Multi-modal Interactions with Language-gesture Enabled (AI) Embodiment (N18-5)
Copied to clipboard
| Challenge: | SMILEE is a conversational agent system that interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures. |
| Approach: | They propose to use a computer-generated avatar to embody a human-machine conversational agent system that interprets verbal utterances and non-verbal behaviors to facilitate natural symmetric multi-modal interactions. |
| Outcome: | The proposed system interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures, and communicates with natural language and gestures through its embodiment as an avatar. |
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
| Approach: | This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
| Outcome: | This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models. |
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (2024.acl-long)
Copied to clipboard
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, Daniel Fried
| Challenge: | Existing benchmarks focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve. |
| Approach: | They propose a benchmark to assess the performance of multimodal web agents . they use visual and textual inputs to process and interpret natural language instructions . |
| Outcome: | a new benchmark assesses the performance of multimodal agents on visually grounded tasks . the benchmark identifies limitations of text-only agents and offers insights towards building stronger agents for the web . |
Connecting Language and Vision to Actions (P18-5)
Copied to clipboard
| Challenge: | Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment. |
| Approach: | This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding. |
| Outcome: | This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog. |
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied. |
| Approach: | They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions. |
| Outcome: | The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions. |
Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations (2025.coling-main)
Copied to clipboard
| Challenge: | Existing data on ESL speakers' communication and interaction skills are lacking in the evaluation of the sophisticated features of dialogue. |
| Approach: | They propose an evaluation framework for interactive dialogue assessment in ESL speakers. |
| Outcome: | The proposed framework provides a means to assess ESL communication, useful for language assessment. |
The VoxWorld Platform for Multimodal Embodied Agents (2022.lrec-1)
Copied to clipboard
| Challenge: | a retrospective of the VoxWorld platform is presented . it is a platform for rapidly building and deploying embodied agents with contextual and situational awareness. |
| Approach: | They present a retrospective on the development of the VoxWorld platform . they focus on three different agent implementations and the functionality needed to accommodate them . |
| Outcome: | The VoxWorld platform has evolved from a theoretical model to a platform capable of multimodal interaction and hybrid reasoning. |
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups. |
| Approach: | They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks. |
| Outcome: | The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants. |
Situated and Interactive Multimodal Conversations (2020.coling-main)
Copied to clipboard
Seungwhan Moon, Satwik Kottur, Paul Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard
| Challenge: | Situated Interactive MultiModal Conversations (SIMMC) is a new direction for virtual assistants that handle multimodal inputs and perform multimodal actions. |
| Approach: | They propose to use Situated Interactive MultiModal Conversations (SIMMC) to train agents to take multimodal actions grounded in a co-evolving multimodal context. |
| Outcome: | The proposed model will be made publicly available. |
SPHERE: An Evaluation Card for Human-AI Systems (2025.findings-acl)
Copied to clipboard
Dora Zhao, Qianou Ma, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang, Tongshuang Wu
| Challenge: | Existing evaluation methods and standards for human-AI systems are unclear, especially for large language models. |
| Approach: | They propose an evaluation card SPHERE which provides a template for evaluation protocols . they outline current evaluation practices and areas for improvement . |
| Outcome: | The evaluation card provides a template for designing evaluation protocols . it outlines current evaluation practices and areas for improvement . |