Papers with MultiAgentBench

    2 papers
    Explicit Trait Inference for Multi-Agent Coordination (2026.acl-long)

    Copied to clipboard

    Challenge: Large language model (LLM) based multi-agent systems (MAS) show promise on complex tasks but remain prone to failures of coordination, such as goal drift, error cascades, and misaligned behaviors.
    Approach: They propose a psychologically grounded method for improving coordination using Explicit Trait Inference (ETI) ETI enables agents to infer and track partner characteristics along two established psychological dimensions—warmth (e.g., trust) and competence (eg. skill)
    Outcome: The proposed method reduces payoff loss in controlled and realistic multi-agent settings by 45–77% and improves performance by 3–29% depending on scenario and model.
    MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents (2025.acl-long)

    Copied to clipboard

    Challenge: Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition.
    Approach: They propose a benchmark to evaluate LLM-based multi-agent systems across diverse, interactive scenarios.
    Outcome: The proposed framework measures task completion and quality of collaboration and competition using novel, milestone-based key performance indicators.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations