Challenge: Multiparty social dialogue is difficult to formalize and expensive to evaluate, especially at scale.
Approach: They propose a lightweight and controllable multi-party dialoguegeneration framework as an experimental instrument for studying generation and evaluation in social interaction.
Outcome: The proposed framework shows that human judgments against state-of-the-art LLM judges are consistent with human preferences for naturalness, engagingness, and overall quality in multi-party social dialogue.

Similar Papers

Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing "LLM-as-a-judge" evaluation frameworks are limited by persona descriptions and are not generalizable to other tasks.
Approach: They propose a framework that can automatically construct multiple evaluator personas with distinct dimensions from relevant text documents and instantiate LLM agents with the persona.
Outcome: The proposed framework can believably simulate human evaluators . it extracts stakeholders' diverse perspectives from the provided research papers and constructs personas for the agents .
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judge (2025.findings-emnlp)

Copied to clipboard

Challenge: LLM-as-Judge frameworks provide scalable alternative to human evaluation . but the question of how intrinsic biases manifest in these settings remains unexplored .
Approach: They conduct systematic analysis of four bias types in multi-agent LLM-as-Judge frameworks . they find debate framework amplifies biases sharply after initial debate .
Outcome: The proposed frameworks amplify biases after debate and show they are stronger in meta-judge scenarios.
Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation (D19-1)

Copied to clipboard

Challenge: Existing evaluation methods for natural language generation are inadequate . distinguishing machine-generated text is challenging even for human evaluators .
Approach: They compare human-based evaluators with automated evaluation procedures . they find human evaluers do not correlate well with discriminative evalators .
Outcome: The proposed evaluation methods are compared with a dozen state-of-the-art generators for online product reviews.
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues (2024.findings-naacl)

Copied to clipboard

Challenge: Modern Large Language Models (LLMs) facilitate high-quality, multi-turn dialogues with humans, but human-based evaluation of such a capability requires substantial manual effort.
Approach: They propose to evaluate LLMs' ability to emulate human-like, multi-turn conversations using an LLM-centric approach.
Outcome: The proposed model emulates human-like, multi-turn conversations using an LLM-centric approach.
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators (2026.findings-eacl)

Copied to clipboard

Challenge: Existing meta-evaluation benchmarks are static, outdated, and lacking in multilingual coverage.
Approach: They propose a framework for curating more representative open-domain dialogue evaluation benchmarks . they leverage several LLMs to generate user-chatbot multilingual dialogues conditioned on varied seed contexts based on a state-of-the-art LLM .
Outcome: The proposed framework exploits state-of-the-art LLMs to perform multilingual evaluations of open-domain chatbots.
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown their potential to deliver human-like judgments.
Approach: They propose a systematic LLM-based multi-agent framework for advanced LLM as-a-judge MT evaluation that integrates dimension-specific results into a final evaluation judgment.
Outcome: The proposed framework outperforms existing LLM-as-a-judge methods and competes with state-of-the-art automatic metrics even when powered by a suboptimal model like GPT-4o mini.
Humans or LLMs as the Judge? A Study on Judgement Bias (2024.emnlp-main)

Copied to clipboard

Challenge: Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs.
Approach: They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges.
Outcome: The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges.
Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations (2024.emnlp-main)

Copied to clipboard

Challenge: Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs.
Approach: They propose a methodological pipeline to investigate model performance across structural attributes of conversations.
Outcome: The proposed method analyzes the performance of an LLM to classify multi-party conversations . it shows that response selection relies more on the textual content of conversations compared to addressee recognition .
DEBATE: Devil’s Advocate-Based Assessment and Text Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for evaluating the quality of machine-generated texts have a relatively low correlation with human performance.
Approach: They propose an NLG evaluation framework based on multi-agent scoring system augmented with a concept of Devil’s Advocate.
Outcome: The proposed evaluation framework outperforms the previous state-of-the-art methods in two meta-evaluation benchmarks in NLG evaluation, SummEval and TopicalChat.
Is ChatGPT a Good Multi-Party Conversation Solver? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are powerful tools for multi-party conversations, but their capacity to handle multi-parties remains unexplored.
Approach: They propose to evaluate ChatGPT and GPT-4's zero-shot learning capabilities within the context of multi-party conversations (MPCs) they also propose to incorporate MPC structures, encompassing both speaker and addressee architecture.
Outcome: The proposed models perform poorly on a number of MPC tasks while GPT-4 performs well on speaker and addressee architecture.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations