Papers with robots

14 papers
Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction (L18-1)

Copied to clipboard

Challenge: a study investigates the influence of gender stereotypes on trust and likability of humanoid robots . explicit gender and stereotypicality of a task are manipulated to influence robot behavior . future research may look into situational variables that drive stereotypification in robot interaction .
Approach: They investigated the influence of gender stereotypes on trust and likability of robots . they used explicit (name and voice) and implicit (personality) genders to manipulate stereotypical tasks . future research may look into situational variables that drive stereotypization .
Outcome: The findings suggest that gender stereotypes need to be differentiated in robot interaction . the gender and personality characteristics of robots influence trust and likability .
Bring the Apple, Not the Sofa: Impact of Irrelevant Context in Embodied AI Commands on VLA Models (2026.eacl-srw)

Copied to clipboard

Challenge: Embodied AI is undergoing rapid development, with robots increasingly exhibiting practical utility in everyday environments.
Approach: They evaluate the robustness of vision language action models under linguistic perturbations . they categorize irrelevant contexts into two groups according to their length and proximity to robot commands .
Outcome: The proposed model can exhibit relative robustness to random context, with a performance drop within 10%, the authors show . human paraphrases of instructions lead to a drop of nearly 20%, the study shows .
An Annotation Approach for Social and Referential Gaze in Dialogue (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on eye gaze information focus on social functions and how it is used in reference resolution.
Approach: They propose an approach for annotating eye gaze considering its social and referential functions in multi-modal dialogue.
Outcome: The proposed annotation scheme is based on eye gaze behavior cues in human-human dialogues.
Learning Physical Common Sense as Knowledge Graph Completion via BERT Data Augmentation and Constrained Tucker Factorization (2020.emnlp-main)

Copied to clipboard

Challenge: Physical commonsense learning is an essential part of human-robot interaction . existing methods of learning physical commons sense suffer from generalization .
Approach: They propose to use physical commonsense learning as a knowledge graph completion problem to better use latent relationships among training samples.
Outcome: The proposed method outperforms existing methods in the human-robot interaction problem.
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for backchannel prediction relied on turn-based or artificially balanced datasets.
Approach: They propose a method for real-time, continuous backchannel prediction using a fine-tuned Voice Activity Projection model.
Outcome: The proposed method outperforms baseline methods in timing and type prediction tasks in real-world environments.
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond (2025.naacl-long)

Copied to clipboard

Challenge: Afrispeech-Dialog is a benchmark dataset of 50 simulated medical and non-medical African-accented English conversations . a 10%+ performance degradation is found in ASR systems on long-form, accented speech .
Approach: They propose to use a dataset to evaluate automatic speech recognition systems on African-accented conversations.
Outcome: The proposed dataset compares state-of-the-art speech recognition systems on accented conversations with native accents and shows a 10%+ performance degradation.
Dialogue Corpus Construction Considering Modality and Social Relationships in Building Common Ground (2022.lrec-1)

Copied to clipboard

Challenge: Several studies have examined the process of building common ground in text chat, but none have investigated the process in depth.
Approach: They constructed a dialogue corpus to investigate the process of building common ground with a particular focus on the modality of dialogue and the social relationship between workers.
Outcome: The results suggest that adding the modality or developing the relationship between workers speeds up the building of common ground.
Learning Adverbs with Spectral Mixture Kernels (2024.findings-acl)

Copied to clipboard

Challenge: In order for robots to collaborate with humans, it is important to share and understand their experiences through language.
Approach: They propose a hierarchical Dirichlet Process-Spectral Mixture Latent Dirichlets Allocation model which learns the relationship between human motions and adverbs by capturing frequency kernels that represent motion characteristics and shared topics of a given aadverts.
Outcome: The proposed model outperforms representative neural network models in terms of perplexity score and predicts more appropriate adverbs.
HADREB: Human Appraisals and (English) Descriptions of Robot Emotional Behaviors (2022.lrec-1)

Copied to clipboard

Challenge: HADREB datasets explore how humans perceive robot emotional states . emotions are a fundamental part of the human language system and are used as scaffolding for language learning .
Approach: They present a dataset of human appraisals and English descriptions of robot emotional behaviors . they use mistyrobotics mist and digital dream labs cozmo robots to analyze the data .
Outcome: The proposed dataset examines how humans perceive robot emotional states and how they relate to human language.
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: Current Vision-Language Models (VLMs) focus on third-person view videos, neglecting the richness of egocentric perceptual experience.
Approach: They propose to use the Egocentric Video Understanding Dataset (EVUD) to train VLMs on video captioning and question answering tasks specific to egocentric videos.
Outcome: The proposed model outperforms open-source models including strong Socratic models using GPT-4 as a planner by 3.6% and outperformed Claude 3 and Gemini Pro Vision 1.0.
J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have ground referential expressions in language to real-world objects for cooperative action generation.
Approach: They propose a Japanese Conversation dataset for real-world reference resolution that ground referential expressions to visual information observed in egocentric views.
Outcome: The proposed dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and assistant robot at home.
Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot Commands (2025.emnlp-main)

Copied to clipboard

Challenge: Existing symbolic parsers lack flexibility to operate in complex, dynamic environments.
Approach: They propose a framework that combines frame semantics with perceptual grounding to enable robots to interpret commands via multimodal logical forms.
Outcome: The proposed framework produces over 11,000 image-command pairs and lowers the cost of manual parsers.
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus (2024.lrec-main)

Copied to clipboard

Challenge: The corpus contains 89,056 utterances and 310,095 words from 278 dialogues averaging 320 utterrances per dialogue.
Approach: They present the Situated Corpus Of Understanding Transactions, a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration.
Outcome: The Situated Corpus Of Understanding Transactions (SCOUT) contains 89,056 utterances and 310,095 words from 278 dialogues averaging 320 utterrances per dialogue.
On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing vision-language-action models are unsuitable for simulated or physical-world deployments . current methods fail when confronted with inherent real-world dynamic variability.
Approach: They propose a test-time reinforcement learning framework that enables on-the-fly policy adaptation during inference.
Outcome: Empirical results show that the proposed framework improves adaptability, stability and task success in dynamic, previously unseen scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations