Challenge: Design-time safety guarantees for human-centered autonomous systems often break down in open-world deployment due to uncertain human interaction.
Approach: They propose an LLM-based architecture that automatically generates personalized safety plans . by itself, the LLM fares poorly at producing safe usage plans, but coupling it with a safety verifier enables the discovery of safe plans.
Outcome: The proposed architecture generates personalized safety plans that are safe for open-world use . the proposed architecture fares poorly at producing safe usage plans .

Similar Papers

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator (2026.acl-long)

Copied to clipboard

Challenge: SafeAgent improves agent safety through fully automated synthetic data generation.
Approach: They propose a framework that improves agent safety through fully automated synthetic data generation.
Outcome: The proposed framework outperforms closed-source models on two safety benchmarks and one real-world task.
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation (2025.findings-acl)

Copied to clipboard

Challenge: Existing alignment methods focus on reactive feedback, where immediate human perception is leveraged to judge sampled model responses as preference data for post-training.
Approach: They propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment.
Outcome: The proposed framework achieves 20% improvement on existing safety benchmarks and an average win rate exceeding 70% against strong baselines.
NITI: Neural Plan Concretization for Incremental Execution, Bridging and Trigger Inference from Underspecified Human Policies (2026.findings-acl)

Copied to clipboard

Challenge: Using NITI, we examine the performance of a safety-critical automated insulin dosing task with minimal contextualization infence overhead.
Approach: They propose a framework that treats large language models as execution-time concretizers of human intent that incrementally executes abstract policies via verifier-grounded interfaces.
Outcome: The proposed framework outperforms one-shot and chain-of-thought baselines on two structurally distinct embodied domains: a world cubing championship 22 Rubik’s Cube scramble and a safety-critical automated insulin dosing task.
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents.
Approach: They propose to integrate human-provided information, feedback, or control into the agent system to enhance system performance, reliability, and safety.
Outcome: The proposed systems improve system performance, reliability, and safety by integrating human-provided information, feedback, or control into the agent system.
AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails (2025.naacl-long)

Copied to clipboard

Challenge: Existing safety-related content safety models are not well-suited for commercial use.
Approach: They propose a taxonomy that can be used to categorize safety risks . it combines human annotations with a multi-LLM "jury" system to assess safety . they plan to open-source Aegis2.0 data and models to aid in safety guardrailing .
Outcome: The proposed taxonomy can be used to assess the safety of human-LLM interactions . it can be trained on large, non-commercial datasets and is open-source .
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation (2025.acl-industry)

Copied to clipboard

Challenge: Clinical note generation (CNG) tools are being developed to address extended working hours and healthcare provider fatigue.
Approach: They evaluate the reliability of 12 open-weight and proprietary LLMs from Anthropic, Meta, Mistral, and OpenAI in CNG in terms of their ability to generate notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity)
Outcome: The results show that the LLMs generated notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity) overall, Meta’s Llama 70B was the most reliable, followed by Mistral’s Small model.
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety evaluations rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail.
Approach: They propose a framework for systematically evaluating the physical safety of LLMs in embodied decision making.
Outcome: The proposed framework assesses the physical safety of LLMs in embodied decision making.
AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems (2025.emnlp-demos)

Copied to clipboard

Challenge: Large language models (LLMs) are being used for planning in orchestrated multi-agent systems . existing LLMs fall short of human expectations and lack effective mechanisms for users to inspect, understand, and control their behaviors.
Approach: They propose a system supporting human-in-the-loop planning through conversational and graph-based interfaces.
Outcome: AIPOM enables users to transparently inspect, refine, and collaboratively guide LLM-generated plans, significantly enhancing user control and trust in multi-agent workflows.
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks (2026.findings-acl)

Copied to clipboard

Challenge: Existing defense methods are insufficient to address in-context reward hacking (ICRH), where LLMs iteratively optimize their behavior to maximize proxy objectives, resulting in harmful side effects.
Approach: They propose a framework that reduces in-context reward hacking (ICRH) through repeated interactions with the environment.
Outcome: The proposed framework reduces ICRH without model fine-tuning while maintaining task performance.
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on developing LLMs to automate complex planning tasks.
Approach: They propose to provide a comprehensive overview of current LLM planners to fill this gap . they examine performance criteria including completeness, executability, optimality, representation, generalization, and efficiency .
Outcome: The proposed survey examines performance criteria for LLM planners and highlights their strengths and weaknesses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations