| Challenge: | Existing studies on comment ranking use user feedback as a quality measure . however, this type of measurement has two drawbacks: (a) user feedback does not always satisfy the service provider's needs, such as to create a fair place; and (b) user input will be biased by where comments appear in a comment thread. |
| Approach: | They propose to evaluate quality of comments on the basis of constructiveness separately from user feedback. |
| Outcome: | The proposed model improves on 100K+ Japanese comments with constructiveness scores . it shows that C-scores are not always related to positive feedback . |
Similar Papers
The Good, the Bad and the Constructive: Automatically Measuring Peer Review’s Utility for Authors (2025.emnlp-main)
Copied to clipboard
| Challenge: | Providing constructive feedback to authors is a core component of peer review . authors lack guidance on how to improve their review, a problem that is often overlooked . |
| Approach: | They use a RevUtil dataset to benchmark fine-tuned models for assessing review comments . they find that machine-generated reviews generally underperform human reviews on these aspects . |
| Outcome: | The proposed model outperforms closed models on four aspects of review comments . the proposed model achieves agreement levels comparable to and exceeding those of human models . |
Automatic Article Commenting: the Task and Dataset (P18-2)
Copied to clipboard
| Challenge: | Existing methods to make comments on articles are based on human-annotated subsets, but they are not suitable for online forums. |
| Approach: | They propose to use a large-scale Chinese corpus with millions of real comments and a human-annotated subset characterizing the comments’ varying quality to generalize a broad set of popular reference-based metrics. |
| Outcome: | The proposed model incorporates human-annotated subset characterizing the comments’ varying quality and shows that it is more accurate than previous models. |
Modeling and Prediction of Online Product Review Helpfulness: A Survey (P18-1)
Copied to clipboard
| Challenge: | review helpfulness modeling is a task that studies the mechanisms that affect review helpfuliness and attempts to accurately predict it. |
| Approach: | This paper provides an overview of the most relevant work in helpfulness prediction . it discusses the insights gained from said work and provides guidelines for future research . |
| Outcome: | This paper summarizes the most relevant work in helpfulness prediction and understanding in the past decade . it outlines the insights gained from the results and provides guidelines for future research . |
Let’s discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Argumentation is a key competence and an important cultural technique in democratic societies. |
| Approach: | They propose to create domain-specific datasets and methods to assess argument quality. |
| Outcome: | The proposed methods address gaps in the literature and aid future research in the domain. |
Reinforced Product Metadata Selection for Helpfulness Assessment of Customer Reviews (D19-1)
Copied to clipboard
| Challenge: | a helpful review is largely concerned with the metadata of its target product . a selector learns from both the key-value product metadata and one of its reviews to take an action . |
| Approach: | They propose a framework that uses product metadata to assess helpfulness of free-text reviews . they use two real-world datasets from amazon.com and Yelp.com to test the framework . |
| Outcome: | The proposed framework can achieve state-of-the-art performance with substantial improvements . it uses two real-world datasets from Amazon.com and Yelp.com . |
On the Role of Reviewer Expertise in Temporal Review Helpfulness Prediction (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for detecting helpful reviews focus on review text and ignore the two key factors of (1) who post the reviews and (2) when the reviews are posted. |
| Approach: | They propose to integrate reviewer's expertise and temporal dynamics to predict helpfulness for unreliable and cold-start reviews. |
| Outcome: | The proposed model improves on existing models and compares with baselines. |
Top-Rank-Focused Adaptive Vote Collection for the Evaluation of Domain-Specific Semantic Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Embedding-based models are increasingly needed for domain-specific evaluation datasets. |
| Approach: | They propose a protocol for the construction of a relatedness-based evaluation dataset based on adaptive pairwise comparisons and appropriate metrics to evaluate a semantic model via the aforementioned dataset. |
| Outcome: | The proposed protocol is particularly accurate in top-rank evaluation. |
The lack of theory is painful: Modeling Harshness in Peer Review Comments (2022.aacl-main)
Copied to clipboard
| Challenge: | a new study shows that peer-review has a power imbalance, making it fraught for authors . authors argue that a little more effort to remain critical but be constructive would help foster a positive outcome . |
| Approach: | They propose to use a dataset to show peer-review comments' harshness scores . they argue that this moderation could help authors to be more constructive . |
| Outcome: | The proposed dataset shows that it can be used to make peer reviews less hurtful and more welcoming. |
It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Moderation is essential for maintaining and improving the quality of online discussions. |
| Approach: | They annotate a dataset on 13 modes of discussion and use it to generate positive moderation. |
| Outcome: | The proposed model shows that professional moderation generates higher ratings than professional moderated moderation, but prefers professional moderate in pairwise comparison. |
Automatic Argument Quality Assessment - New Datasets and Methods (D19-1)
Copied to clipboard
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
| Challenge: | 6.3k arguments were collected from contributors of various levels, and are released as part of this work. |
| Approach: | They propose to use a language model to annotate arguments for argument ranking and argument-pair classification. |
| Outcome: | The proposed methods outperform state-of-the-art methods in the argument ranking task and argument-pair classification task. |