Challenge: Rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support.
Approach: They propose a peer review benchmarking tool based on paper-specific rubrics and a rubric-guided framework that decomposes reviewing into drafting and grounding stages.
Outcome: The proposed framework outperforms baselines with stronger/larger backbones in both alignment with human judgments and rubric-based review quality across 8 dimensions.

Similar Papers

ReviewEval: An Evaluation Framework for AI-Generated Reviews (2025.findings-emnlp)

Copied to clipboard

Challenge: escalating volume of academic research necessitates innovative approaches to peer review . authors propose reviewEval, ReviewAgent and ReviewEval to improve on existing reviews .
Approach: They propose a framework for AI-generated reviews that measures alignment with human assessments . they propose 'reviewAgent' that iteratively optimizes its intermediate outputs and external improvement loops .
Outcome: The proposed framework improves actionable insights and analytical depth by 6.78% and 47.62% over baselines and expert reviews.
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring (2025.findings-emnlp)

Copied to clipboard

Challenge: a new framework for automated essay scoring is needed to achieve multi-perspective understanding and judgment.
Approach: They propose a roundtable essay scoring framework that performs precise and human-aligned scoring under a zero-shot setting.
Outcome: The proposed framework outperforms previous zero-shot AES approaches by enabling collaboration among agents with diverse evaluation perspectives.
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic (2026.findings-acl)

Copied to clipboard

Challenge: Existing search agent pipelines rely on sparse outcome rewards, leading to inefficient exploration and unstable training.
Approach: They propose a tool-integrated reasoning framework that provides turn-level feedback via a retrospective critic mechanism.
Outcome: The proposed framework outperforms baselines in multi-hop reasoning benchmarks and achieves faster convergence and training stability.
Preference Optimization for Review Question Generation Improves Writing Quality (2026.findings-acl)

Copied to clipboard

Challenge: Peer reviewers are overloaded and face tight deadlines, leading some to rely on large language models (LLMs) to draft questions and comments.
Approach: They use open-review review datasets to train a human preference model based on human reviewer questions . human evaluations show IntelliAsk generates more grounded, substantive and effortful questions than strong baselines .
Outcome: The proposed model predicts reviewer-question quality better than API-based SFT baselines and provides scalable evaluation.
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality (2026.findings-acl)

Copied to clipboard

Challenge: Increasing use of large language models (LLMs) in academic review has raised concerns about quality and fairness.
Approach: They propose a framework to improve the quality of LLM-generated reviews by using retrieval-augmented generation.
Outcome: The proposed framework improves the human-level quality of LLM-generated reviews by adopting prompt engineering and retrieval-augmented generation.
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process (2025.acl-long)

Copied to clipboard

Challenge: Existing Large Language Models (LLMs) face limited domain expertise, hallucinated reasoning, and a lack of structured evaluation.
Approach: They propose a multi-stage framework to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation.
Outcome: The proposed model outperforms CycleReviewer-70B with fewer tokens and achieves 88.21% and 80.20% win rates.
IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review (2026.acl-long)

Copied to clipboard

Challenge: Scientific research relies on accurate information retrieval from literature to support analytical decisions.
Approach: They propose a task that automates fine-grained information retrieval *faithfully* grounded in the provided content in response to research-driven queries.
Outcome: The proposed agent achieves 13.2% higher cross-domain accuracy than state-of-the-art RAG and research-agent baselines across seven backbone LLMs.
AnalystBench: Benchmarking professional long-form report generation with web-mined multimodal tasks (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks decompose the end-to-end professional report generation into individual components.
Approach: They propose a benchmarking tool that evaluates 20 real-world professional report generation tasks grounded in multimodal document collections.
Outcome: The proposed model outperforms closed-source models on executive summarization tasks but drops significantly on long-horizon synthesis tasks.
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can automatically draft reviews, but determining whether they are trustworthy requires systematic evaluation.
Approach: They propose an automatic focus-level evaluation pipeline based on two sets of facets . authors evaluated LLM reviews at surface-level or content-level .
Outcome: The proposed framework enables automatic evaluation of paper reviews based on two sets of facets . the framework compared open review paper reviews with human experts on validity, clarity, novelty .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations