Papers with MAD

15 papers
SELENE: Selective and Evidence-Weighted LLM Debating for Efficient and Reliable Reasoning (2026.eacl-industry)

Copied to clipboard

Challenge: Existing multi-agent debate frameworks are computationally expensive and prone to degradation under pro-longed debates due to redundant exchanges and unstable judging.
Approach: They propose a framework that unifies Selective Debate Initiation (SDI) with Evidence Weighted Self-Consistency (EWSC) for adaptive, debate-on-demand reasoning.
Outcome: Evaluated on BoolQ, CosmosQA, and an internal QnA benchmark, the proposed framework achieves higher factual robustness and efficiency.
MALLM: Multi-Agent Large Language Models Framework (2025.emnlp-demos)

Copied to clipboard

Challenge: Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise.
Approach: They propose an open-source framework that enables systematic analysis of multi-agent debates.
Outcome: The proposed framework enables systematic analysis of multi-agent debate components.
Enhancing Multi-Agent Debate System Performance via Confidence Expression (2025.findings-emnlp)

Copied to clipboard

Challenge: Multi-Agent Debate systems leverage multiple LLMs to improve task performance.
Approach: They propose to integrate confidence expression into MAD systems to help LLMs communicate their confidence levels.
Outcome: The proposed approach improves debate effectiveness and overall system performance by integrating confidence expression into MAD systems.
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding (2023.acl-long)

Copied to clipboard

Challenge: Existing work on video temporal grounding for long videos is limited by existing datasets.
Approach: They propose a query-guided window selection strategy and a coarse-to-fine mechanism to speed up inference for long videos.
Outcome: The proposed framework accelerates inference time by 2x on Ego4D-NLQ and 15x on MAD while keeping SOTA results.
S2-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency (2025.naacl-long)

Copied to clipboard

Challenge: Large language models exhibit limitations when handling complex mathematical reasoning and logical inference tasks.
Approach: They propose a sparsification strategy to reduce token costs within Multi-agent Debate (MAD) this strategy minimizes ineffective exchanges of information and unproductive discussions among agents .
Outcome: The proposed approach reduces token costs by up to 94.5% while maintaining performance degradation below 2.0%.
Explain-Analyze-Generate: A Sequential Multi-Agent Collaboration Method for Complex Reasoning (2025.coling-main)

Copied to clipboard

Challenge: Multiagent debate (MAD) is a popular approach for large language models . however, the performance of LLMs is suboptimal in complex reasoning scenarios .
Approach: They propose a sequential collaboration framework to enable agents to provide constructive assistance to peers by decomposing complex tasks into essential subtasks.
Outcome: The proposed framework achieves the highest performance on 19 out of 23 tasks and lower costs compared to MAD.
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve the reasoning performance of LLMs suffer from two major shortcomings: too lengthy input contexts and overconfidence dilemma.
Approach: They propose a method to debating among LLM agents using a sparse debator graph . they use a module called McKinsey-based Debate Matter to optimize the debators .
Outcome: The proposed method has been well demonstrated across eight datasets from four task types.
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluation of large language models are inefficient and inefficient due to inaccuracy of standard metrics in human perception of text quality and inefficiency in sampling informative test examples.
Approach: They propose a sample-efficient human evaluation method for large language models based on the principle of MAximum Discrepancy (MAD) competition.
Outcome: The proposed method achieves the “golden” ranking of LLMs with a minimum set of input instructions, which in turn reveal their relative strengths and weaknesses.
When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Multi-agent debate (MAD) aims to improve large language model reasoning by letting multiple agents exchange answers and then aggregate their opinions.
Approach: They propose a principled framework that joins sycophancy and self-bias to mitigate and quantify identity bias in multi-agent debate by removing identity markers from prompts.
Outcome: The proposed framework joins identity-driven sycophancy and self-bias to mitigate and quantify identity bias in multi-agent debate.
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Despite advances in large language models, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification.
Approach: They propose a multi-agent debate framework that reformulates misinformation detection as a structured adversarial debate based on fact-checking workflows .
Outcome: The proposed framework enables iterative refinement of evidence while improving decision transparency.
Removal of Hallucination on Hallucination: Debate-Augmented RAG (2025.acl-long)

Copied to clipboard

Challenge: erroneous or biased retrieval can mislead generation, compounding hallucinations.
Approach: They propose a framework that integrates multi-agent debates into retrieval and generation stages to improve retrieval reliability.
Outcome: The proposed framework improves retrieval reliability, reduces hallucinations and significantly improves overall factual accuracy.
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate (2024.emnlp-main)

Copied to clipboard

Challenge: Modern large language models (LLMs) have shown remarkable performance on general language tasks but struggle on complex reasoning tasks.
Approach: They propose a multi-agent debate framework that encourages divergent thinking in LLMs . they propose to break debate and use a judge to obtain a final solution .
Outcome: The proposed framework encourages divergent thinking in large language models . it is able to generate novel thoughts even if initial stance is incorrect .
The Effects of Pretraining in Video-Guided Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to improve VMT models integrate text and video modalities.
Approach: They propose an approach that improves the performance of VMT models by using a new dataset which contains transcribed audio descriptions of movies.
Outcome: The proposed model improves on the MAD (Movie Audio Descriptions) dataset.
Free-MAD: Consensus-Free Multi-Agent Debate (2026.findings-acl)

Copied to clipboard

Challenge: Existing multi-agent debate methods rely on multiple rounds of interaction among agents to reach consensus, and the final output is decided by majority voting in the last round.
Approach: They propose a multi-agent debate framework that eliminates the need for consensus among agents and reconstructs the debate phase by introducing anti-conformity.
Outcome: Experiments on eight benchmark datasets show that Free-MAD significantly improves reasoning performance while requiring only a single-round debate and thus reducing token costs.
Demystifying Multi-Agent Debate: The Role of Confidence and Diversity (2026.findings-acl)

Copied to clipboard

Challenge: Multi-agent debate (MAD) is widely used to improve large language models' (LLMs) reasoning and test-time scaling.
Approach: They propose a diversity-aware initialisation that selects a more diverse pool of candidate answers, increasing the likelihood that a correct hypothesis is present at the start of debate.
Outcome: The proposed protocol outperforms vanilla MAD and majority vote on six reasoning-oriented QA benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations