Challenge: Many group decisions are open-ended, and aggregation approaches suppress minority perspectives . team members must surface hidden assumptions, discuss disagreements, negotiate acceptable trade-offs .
Approach: They propose a multi-agent system that instantiates a proxy agent for each team member . they also conduct a structured discussion to elicit agreements and disagreements .
Outcome: The proposed system outperforms direct aggregation on two teamwork tasks . it can judge how well individual views are represented in team decisions and consensually good deliverables .

Similar Papers

Multi-Agent Collaboration via Cross-Team Orchestration (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have significantly impacted various domains, especially through organized LLM-driven autonomous agents.
Approach: They propose a framework that enables orchestrated teams to jointly propose various task-oriented solutions and interact with their insights in a self-independence while cross-team collaboration environment for superior solutions generation.
Outcome: Experiments show that the framework can generate better software quality compared to state-of-the-art frameworks.
LLM Agents for Coordinating Multi-User Information Gathering (2025.findings-acl)

Copied to clipboard

Challenge: Recent large language models (LLMs) are becoming a crucial building block in developing automated agents that can assist human users with complex tasks.
Approach: They introduce PeopleJoin, a benchmark for evaluating LM-mediated collaborative problem solving.
Outcome: The proposed benchmarks are adapted from existing benchmarks for database question answering and multi-document summarization.
Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing frameworks prioritize structural architectures and role assignments but neglect granular mechanics of agent collaboration.
Approach: They propose to use centralized governance, instructor-led participation, ordered interaction patterns to optimize task accuracy and computational efficiency.
Outcome: The proposed model improves task accuracy and computational efficiency under two context-dependent scenarios.
DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing data insight agents fail to deliver satisfactory results due to insufficient utilization of domain knowledge, shallow analytical depth, and error-prone code generation.
Approach: They propose a novel multi-agent framework that incorporates external knowledge retrieval to enrich the analytical context, a multi-role debating mechanism to simulate diverse analytical perspectives and deepen analytical depth, and multi-path reasoning to improve the accuracy of the generated code and insights.
Outcome: Extensive experiments on InsightBench show that DataSage outperforms existing data insight agents across all difficulty levels, improving by 7.5% and 13.9% respectively in insight-level and summary-level metrics.
MAPLE: Multi-Aspect Panels of LLM Evaluators for Open-Ended Questions (2026.findings-acl)

Copied to clipboard

Challenge: LLM-as-a-Judge uses LLMs to evaluate open-ended questions . however, the discrepancy between LLM generated evaluations and human evaluations remains a critical problem in this field .
Approach: They propose a framework that orchestrates evaluations across multiple criteria using multiple LLMs.
Outcome: The proposed framework achieves superior alignment with human evaluations compared to baselines.
WorkTeam: Constructing Workflows from Natural Language with Multi-Agents (2025.naacl-industry)

Copied to clipboard

Challenge: Existing workflow construction methods require specialized knowledge and task-switching skills.
Approach: They propose a multi-agent workflow framework that incorporates a supervisor, orchestrator, and filler agent.
Outcome: The proposed framework significantly increases the success rate of workflow construction . the proposed framework is based on a dataset of 3,695 real-world business samples .
Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key? (2024.acl-long)

Copied to clipboard

Challenge: Recent progress in LLMs discussion suggests that multi-agent discussion improves the reasoning abilities of LLM.
Approach: They propose a group discussion framework to enrich the set of discussion mechanisms.
Outcome: The proposed framework performs better on a wide range of reasoning tasks and backbone LLMs.
CoopValue: Revealing LLM Value Preferences Through Multi-Agent Cooperation (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of large language models rely on single-agent dilemmas or static binary-choice tasks, offering limited insight into how cooperation contexts influence LLM behavior.
Approach: They propose a multi-agent evaluation framework that assesses LLMs’ value preferences through cooperative scenarios.
Outcome: The proposed framework assesses LLMs’ value preferences through cooperative scenarios.
CrowdAgent: Multi-Agent Managed Multi-Source Annotation System (2025.emnlp-demos)

Copied to clipboard

Challenge: Recent approaches to annotate data focus on labeling, but lack holistic process control . a novel system that integrates task assignment, data annotation, and quality/cost management is needed .
Approach: They propose a multi-agent system that integrates task assignment, data annotation, and quality/cost management.
Outcome: The proposed system automates human management by using a collaborative multi-agent system.
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing "LLM-as-a-judge" evaluation frameworks are limited by persona descriptions and are not generalizable to other tasks.
Approach: They propose a framework that can automatically construct multiple evaluator personas with distinct dimensions from relevant text documents and instantiate LLM agents with the persona.
Outcome: The proposed framework can believably simulate human evaluators . it extracts stakeholders' diverse perspectives from the provided research papers and constructs personas for the agents .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations