An Evaluation Framework for Multimodal Interaction (L18-1)

Copied to clipboard

Challenge: a framework for evaluating multimodal interactions is presented . it leverages the semantics of language and gesture to assess mutual understanding . consistent evaluation is required to test areas where the system needs improvement .
Approach: They propose a framework for evaluating interactions between human and virtual agent . they use VoxML as a platform to model interactions using natural language and gesture .
Outcome: The proposed framework assesses the level of mutual understanding and ease of communication between human and computer agents in a blocks world scenario.

Similar Papers

SMILEE: Symmetric Multi-modal Interactions with Language-gesture Enabled (AI) Embodiment (N18-5)

Copied to clipboard

Challenge: SMILEE is a conversational agent system that interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures.
Approach: They propose to use a computer-generated avatar to embody a human-machine conversational agent system that interprets verbal utterances and non-verbal behaviors to facilitate natural symmetric multi-modal interactions.
Outcome: The proposed system interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures, and communicates with natural language and gestures through its embodiment as an avatar.
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Approach: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Outcome: This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models.
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve.
Approach: They propose a benchmark to assess the performance of multimodal web agents . they use visual and textual inputs to process and interpret natural language instructions .
Outcome: a new benchmark assesses the performance of multimodal agents on visually grounded tasks . the benchmark identifies limitations of text-only agents and offers insights towards building stronger agents for the web .
Connecting Language and Vision to Actions (P18-5)

Copied to clipboard

Challenge: Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment.
Approach: This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding.
Outcome: This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog.
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied.
Approach: They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions.
Outcome: The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions.
Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing data on ESL speakers' communication and interaction skills are lacking in the evaluation of the sophisticated features of dialogue.
Approach: They propose an evaluation framework for interactive dialogue assessment in ESL speakers.
Outcome: The proposed framework provides a means to assess ESL communication, useful for language assessment.
The VoxWorld Platform for Multimodal Embodied Agents (2022.lrec-1)

Copied to clipboard

Challenge: a retrospective of the VoxWorld platform is presented . it is a platform for rapidly building and deploying embodied agents with contextual and situational awareness.
Approach: They present a retrospective on the development of the VoxWorld platform . they focus on three different agent implementations and the functionality needed to accommodate them .
Outcome: The VoxWorld platform has evolved from a theoretical model to a platform capable of multimodal interaction and hybrid reasoning.
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
Situated and Interactive Multimodal Conversations (2020.coling-main)

Copied to clipboard

Challenge: Situated Interactive MultiModal Conversations (SIMMC) is a new direction for virtual assistants that handle multimodal inputs and perform multimodal actions.
Approach: They propose to use Situated Interactive MultiModal Conversations (SIMMC) to train agents to take multimodal actions grounded in a co-evolving multimodal context.
Outcome: The proposed model will be made publicly available.
SPHERE: An Evaluation Card for Human-AI Systems (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods and standards for human-AI systems are unclear, especially for large language models.
Approach: They propose an evaluation card SPHERE which provides a template for evaluation protocols . they outline current evaluation practices and areas for improvement .
Outcome: The evaluation card provides a template for designing evaluation protocols . it outlines current evaluation practices and areas for improvement .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations