Papers with VirtualHome

3 papers
VestaBench: An Embodied Benchmark for Safe Long-Horizon Planning Under Multi-Constraint and Adversarial Settings (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing safety benchmarks do not represent a diverse range of multi-constraint tasks that require long-horizon planning with a focus on safety.
Approach: They propose a benchmark to assess the safety of embodied AI agents under multiple constraints.
Outcome: The proposed benchmarks show that LLMs perform poorly against their tasks . they also suffer significantly compromised safety outcomes .
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language planning methods struggle with long-horizon reasoning in dynamic environments due to the difficulty of training models to generate high-quality reasoning processes.
Approach: They propose a framework that enhances reasoning and action selection for long-horizon task planning through structured evaluation and optimized training.
Outcome: The proposed framework outperforms existing methods on short-horizon tasks but struggles with long-horizon reasoning in dynamic environments.
OntoGuard: Enforcing Action Admissibility for LLM Agents in Complex Interactive Environments (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to large language models (LLMs) are limited by their ability to enforce environmental and behavioral admissibility.
Approach: They propose an ontological framework to guard LLM agents by enforcing environmental and behavioral admissibility.
Outcome: Experiments on ScienceWorld and VirtualHome show that OntoGuard can enforce environmental and behavioral admissibility while preventing invalid actions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations