Challenge: despite its importance, little work exists on modelling rigour in scientific writing . despite widespread use of term, scientific literature lacks definition of rigor .
Approach: They propose a framework to automatically identify and define rigour criteria and assess their relevance in scientific writing.
Outcome: The proposed framework can be tailored to the evaluation of scientific rigour for different areas.

Similar Papers

RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support.
Approach: They propose a multimodal framework that retrieves supporting evidence from a paper and assigns each claim an overstatement score.
Outcome: The proposed framework retrieves supporting evidence from ICLR and NeurIPS papers and assigns each claim an overstatement score.
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches.
Approach: They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries.
Outcome: The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Automatic Compilation of Resources for Academic Writing and Evaluating with Informal Word Identification and Paraphrasing System (2020.lrec-1)

Copied to clipboard

Challenge: a systematic review of academic writing aids aims to build a writing aid system that automatically edits a text to adhere to the academic style of writing.
Approach: They propose to build a writing aid system that automatically edits a text to adhere to the academic style of writing.
Outcome: The proposed system outperforms existing academic resources in terms of word identification and ranking . the informal word identification component achieves an F-1 score of 82% .
MiST: a Large-Scale Annotated Resource and Neural Models for Functions of Modal Verbs in English Scientific Text (2022.findings-emnlp)

Copied to clipboard

Challenge: modal verbs are used for hedges, but they may also denote abilities and restrictions in scientific texts . modals are often used for hedging, but prior work on this topic has been limited .
Approach: They propose a dataset that contains 3737 modal instances in five scientific domains . they evaluate a set of competitive neural architectures to model the distinctions in MIST .
Outcome: The proposed dataset contains 3737 modal instances in five scientific domains . leveraging non-scientific data is of limited benefit for modeling the distinctions in MIST .
Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm.
Approach: They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
Outcome: The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language model reasoning focus on mathematics and coding domains, but scientific reasoning remains limited in other domains due to limited dataset coverage.
Approach: They propose a framework for sustainable scientific reasoning QA generation by synthesizing a new dataset of domain-specific science questions from peer-reviewed literature.
Outcome: The proposed framework and dataset enable scalable and sustainable research in scientific reasoning.
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context.
Approach: They propose a task of context-aware text generation in the scientific domain to exploit the contributions of context in generated texts.
Outcome: The proposed dataset comprehensively benchmarks the efficacy of the proposed dataset in generating description and paragraph.
arXivEdits: Understanding the Human Revision Process in Scientific Writing (2022.emnlp-main)

Copied to clipboard

Challenge: a new computational framework is developed to study text revision in scientific writing . authors propose a method to extract revision at document-, sentence-, and word-levels .
Approach: They propose a computational framework for studying text revision in scientific writing . arXivEdits is an annotated corpus of 751 full papers from arX . authors propose to use sentence alignment, fine-grained edits and intents to extract revision .
Outcome: The proposed framework can be used to study revision in scientific writing.
Writing Strategies for Science Communication: Data and Computational Analysis (2020.emnlp-main)

Copied to clipboard

Challenge: Existing science communication guides do not provide empirical evidence for how their strategies are used in practice.
Approach: They propose to use prescriptive writing strategies to identify and train human-readable annotations that can be automatically recognized by a corpus of 128k science writing documents in English.
Outcome: The proposed system can be used to detect and suggest writing strategies for scientists by allowing them to automatically recognize them.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations