On the Rigour of Scientific Writing: Criteria, Analysis, and Insights (2024.findings-emnlp)
Copied to clipboard
| Challenge: | despite its importance, little work exists on modelling rigour in scientific writing . despite widespread use of term, scientific literature lacks definition of rigor . |
| Approach: | They propose a framework to automatically identify and define rigour criteria and assess their relevance in scientific writing. |
| Outcome: | The proposed framework can be tailored to the evaluation of scientific rigour for different areas. |
Similar Papers
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. |
| Approach: | They propose a multimodal framework that retrieves supporting evidence from a paper and assigns each claim an overstatement score. |
| Outcome: | The proposed framework retrieves supporting evidence from ICLR and NeurIPS papers and assigns each claim an overstatement score. |
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches. |
| Approach: | They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries. |
| Outcome: | The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Automatic Compilation of Resources for Academic Writing and Evaluating with Informal Word Identification and Paraphrasing System (2020.lrec-1)
Copied to clipboard
| Challenge: | a systematic review of academic writing aids aims to build a writing aid system that automatically edits a text to adhere to the academic style of writing. |
| Approach: | They propose to build a writing aid system that automatically edits a text to adhere to the academic style of writing. |
| Outcome: | The proposed system outperforms existing academic resources in terms of word identification and ranking . the informal word identification component achieves an F-1 score of 82% . |
MiST: a Large-Scale Annotated Resource and Neural Models for Functions of Modal Verbs in English Scientific Text (2022.findings-emnlp)
Copied to clipboard
| Challenge: | modal verbs are used for hedges, but they may also denote abilities and restrictions in scientific texts . modals are often used for hedging, but prior work on this topic has been limited . |
| Approach: | They propose a dataset that contains 3737 modal instances in five scientific domains . they evaluate a set of competitive neural architectures to model the distinctions in MIST . |
| Outcome: | The proposed dataset contains 3737 modal instances in five scientific domains . leveraging non-scientific data is of limited benefit for modeling the distinctions in MIST . |
Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)
Copied to clipboard
| Challenge: | Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm. |
| Approach: | They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning. |
| Outcome: | The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning. |
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language model reasoning focus on mathematics and coding domains, but scientific reasoning remains limited in other domains due to limited dataset coverage. |
| Approach: | They propose a framework for sustainable scientific reasoning QA generation by synthesizing a new dataset of domain-specific science questions from peer-reviewed literature. |
| Outcome: | The proposed framework and dataset enable scalable and sustainable research in scientific reasoning. |
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context. |
| Approach: | They propose a task of context-aware text generation in the scientific domain to exploit the contributions of context in generated texts. |
| Outcome: | The proposed dataset comprehensively benchmarks the efficacy of the proposed dataset in generating description and paragraph. |
arXivEdits: Understanding the Human Revision Process in Scientific Writing (2022.emnlp-main)
Copied to clipboard
| Challenge: | a new computational framework is developed to study text revision in scientific writing . authors propose a method to extract revision at document-, sentence-, and word-levels . |
| Approach: | They propose a computational framework for studying text revision in scientific writing . arXivEdits is an annotated corpus of 751 full papers from arX . authors propose to use sentence alignment, fine-grained edits and intents to extract revision . |
| Outcome: | The proposed framework can be used to study revision in scientific writing. |
Writing Strategies for Science Communication: Data and Computational Analysis (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing science communication guides do not provide empirical evidence for how their strategies are used in practice. |
| Approach: | They propose to use prescriptive writing strategies to identify and train human-readable annotations that can be automatically recognized by a corpus of 128k science writing documents in English. |
| Outcome: | The proposed system can be used to detect and suggest writing strategies for scientists by allowing them to automatically recognize them. |