Automatic Analysis of Substantiation in Scientific Peer Reviews (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing systems to analyze peer reviews' quality are inadequate due to the increasing workload of reviewers and the lack of domain experts . |
| Approach: | They propose to use a claim-evidence pair extraction problem to analyze substantiation in peer reviews and train an argument mining system to do the same. |
| Outcome: | The proposed system could be used by conference managers and reviewers to analyze the quality of peer reviews. |
Similar Papers
Argument Mining for Understanding Peer Reviews (N19-1)
Copied to clipboard
| Challenge: | In 2015 alone, approximately 63.4 million hours were spent on peer reviews. |
| Approach: | They propose to automatically detect argumentative propositions put forward by reviewers and their types by automatically detecting their types and types. |
| Outcome: | The proposed method detects (1) the argumentative propositions put forward by reviewers, and (2) their types (e.g., evaluating the work or making suggestions for improvement). |
Automatic Argument Quality Assessment - New Datasets and Methods (D19-1)
Copied to clipboard
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
| Challenge: | 6.3k arguments were collected from contributors of various levels, and are released as part of this work. |
| Approach: | They propose to use a language model to annotate arguments for argument ranking and argument-pair classification. |
| Outcome: | The proposed methods outperform state-of-the-art methods in the argument ranking task and argument-pair classification task. |
Seeing Things from a Different Angle:Discovering Diverse Perspectives about Claims (N19-1)
Copied to clipboard
| Challenge: | a number of fact checking techniques are used to identify and eliminate biases in text data. |
| Approach: | They propose to use search engines to expand and diversify a dataset of claims, perspectives and evidence to address a selection bias. |
| Outcome: | The proposed approach outperforms existing methods in a language understanding task. |
Is Something Better than Nothing? Automatically Predicting Stance-based Arguments Using Deep Learning and Small Labelled Dataset (N18-2)
Copied to clipboard
| Challenge: | Argument mining is a subset of NLP that deals with extracting arguments from user-based content. |
| Approach: | They propose to use weakly supervised and semi-supervised methods to automatically annotate reviews and provide large annotated datasets. |
| Outcome: | The proposed methods can be used to learn better models for implicit/explicit opinion classification. |
Towards a Perspectivist Turn in Argument Quality Assessment (2025.naacl-long)
Copied to clipboard
| Challenge: | Argument quality is a key aspect of computational argumentation (CA), but it still exhibits a high degree of subjectivity in perception. |
| Approach: | They propose to use a multi-layered classification to target two aspects of argument quality in a systematic review of NLP datasets. |
| Outcome: | The proposed model improves the quality of annotators and their ability to be used in perspectivist research. |
A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications (N18-1)
Copied to clipboard
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, Roy Schwartz
| Challenge: | a dataset of 14.7K paper drafts and accept/reject decisions in top-tier venues including ACL, NIPS and ICLR is presented to study peer reviews. |
| Approach: | They propose to use the dataset to collect peer reviews from top-tier venues including ACL, NIPS and ICLR and to use it to create a dataset of peer reviews for research purposes. |
| Outcome: | The proposed dataset includes 14.7K paper drafts and accept/reject decisions in top-tier venues including ACL, NIPS and ICLR. |
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches frame reviewer disagreement as binary contradiction detection over isolated sentence pairs, abstracting away review context and obscuring differences in severity of evaluative conflict. |
| Approach: | They propose a fine-grained formulation of reviewer contradiction analysis that operates over full peer reviews by explicitly identifying contradiction evidence spans and assigning graded disagreement intensity scores. |
| Outcome: | The proposed framework outperforms strong single-agent and generic multi-agend baselines in evidence identification and intensity agreement. |
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework (2026.tacl-1)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used as fully automatic review generators (ARGs). |
| Approach: | They propose a fully automated counterfactual evaluation framework that isolates and tests a core review skill that underpins high-quality peer review: detecting faulty research logic. |
| Outcome: | The proposed framework isolates and tests a range of ARG approaches and shows that flaws in research logic have no significant effect on their output reviews. |
APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task Learning (2020.emnlp-main)
Copied to clipboard
| Challenge: | Argument mining is an important research field that attracts growing attention in recent years. |
| Approach: | They propose a new task to extract argument pairs from peer review and rebuttal . they use an open review platform to analyze the contents, structure and connections . |
| Outcome: | The proposed task is based on a dataset of 4,764 fully annotated review-rebuttal passage pairs . it is able to detect argumentative propositions and extract argument pairs from the corpus . |
DeepSentiPeer: Harnessing Sentiment in Review Texts to Recommend Peer Review Decisions (P19-1)
Copied to clipboard
| Challenge: | Existing peer review system is not straightforward and requires domain knowledge, expertise, and intelligence of human reviewers, which is somewhat elusive with the current state of AI. |
| Approach: | They propose to use peer review texts to predict acceptance or rejection of a manuscript based on reviewer sentiment. |
| Outcome: | The proposed deep neural architecture achieves significant performance improvement over baselines (29% error reduction) in a recently released dataset of peer reviews. |