Challenge: Recent studies have focused on generic AI-generated text detection or estimating fraction of peer-reviews that can be AI-generated.
Approach: They propose a model that detects whether a peer-review is written by ChatGPT and a reviewer-generated model that generates similar outputs upon re-prompting.
Outcome: The proposed model is more robust, but paraphrasing is more effective.

Similar Papers

MixRevDetect: Towards Detecting AI-Generated Content in Hybrid Peer Reviews. (2025.naacl-short)

Copied to clipboard

Challenge: Existing methods for detecting fully AI-generated peer reviews fail to detect finer-grained AI-generated points within mixed-authorship reviews.
Approach: They propose a method to identify AI-generated points in peer reviews using large language models . their approach achieved an F1 score of 88.86%, significantly outperforming existing methods .
Outcome: The proposed method outperforms existing methods in identifying AI-generated points in peer reviews.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)

Copied to clipboard

Challenge: Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity).
Approach: They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions.
Outcome: The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization.
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) motivated methods that assist or automate different stages of peer review pipeline.
Approach: They synthesize techniques to enhance peer review generation and after-review tasks aligned to reviews.
Outcome: The proposed methods improve the peer review process by fine-tuning strategies, agent-based systems, and emerging paradigms.
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance .
Approach: They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy.
Outcome: The proposed model will be able to detect human-written content in real time.
A Practical Examination of AI-Generated Text Detectors for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to detect large language models are prone to misuse, such as generating fake news articles, facilitating academic plagiarism or spamming.
Approach: They evaluate several popular detectors to evaluate their effectiveness against a range of domains, datasets, and models.
Outcome: The proposed methods perform poorly in certain settings, with TPR@.01 as low as 0%.
ReviewEval: An Evaluation Framework for AI-Generated Reviews (2025.findings-emnlp)

Copied to clipboard

Challenge: escalating volume of academic research necessitates innovative approaches to peer review . authors propose reviewEval, ReviewAgent and ReviewEval to improve on existing reviews .
Approach: They propose a framework for AI-generated reviews that measures alignment with human assessments . they propose 'reviewAgent' that iteratively optimizes its intermediate outputs and external improvement loops .
Outcome: The proposed framework improves actionable insights and analytical depth by 6.78% and 47.62% over baselines and expert reviews.
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media (2025.acl-long)

Copied to clipboard

Challenge: Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs) however, the misuse of AIGTs could have profound implications for public opinion .
Approach: They collect a dataset with 2.4M posts from 3 major social media platforms . they then construct a diverse dataset to train and evaluate AIGT detectors .
Outcome: The proposed dataset analyzes 2.4M posts from 3 major social media platforms from 2022 to 2024 . it finds that Medium and Quora show marked increases in AAR .
Generative Reviewer Agents: Scalable Simulacra of Peer Review (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing peer review mechanisms are limited by the small fraction of researchers with established networks.
Approach: They propose a system that extends a large language model and equips agents with reviewer personas derived from historical data to enable generative reviewers.
Outcome: The proposed architecture performs comparable to human reviewers in providing detailed feedback and predicting paper outcomes.
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? (2026.acl-long)

Copied to clipboard

Challenge: Existing evidence suggests that LLMs are not able to detect scientifically unsound work from malicious or poorly designed research agents.
Approach: They develop a framework that evaluates whether fabrication-oriented paper generation agents can deceive multi-model LLM review systems.
Outcome: The proposed framework shows that fabricated papers achieve acceptance rates up to 18% . the framework shows only marginal improvements, with detection accuracy barely exceeding random chance.
How Reliable Are AI-Generated-Text Detectors? An Assessment Framework Using Evasive Soft Prompts (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect AI-generated text are inadequate, causing misuse of the text.
Approach: They propose a universal evasive prompt framework that can prompt any PLM to generate “human-like” text that can mislead detectors.
Outcome: The proposed approach can prompt any PLM to generate “human-like” text that can mislead detectors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations