Challenge: Existing science communication guides do not provide empirical evidence for how their strategies are used in practice.
Approach: They propose to use prescriptive writing strategies to identify and train human-readable annotations that can be automatically recognized by a corpus of 128k science writing documents in English.
Outcome: The proposed system can be used to detect and suggest writing strategies for scientists by allowing them to automatically recognize them.

Similar Papers

Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm.
Approach: They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
Outcome: The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
AI for Science in the Era of Large Language Models (2024.emnlp-tutorials)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated significant prowess in tasks involving natural language, such as translating languages, constructing chatbots, and answering questions.
Approach: This tutorial explores the application of large language models to three crucial categories of scientific data: 1) textual data, 2) biomedical sequences, and 3) brain signals.
Outcome: This tutorial will explore the application of large language models to three crucial categories of scientific data.
Generating Scientific Definitions with Controllable Complexity (2022.acl-long)

Copied to clipboard

Challenge: Unfamiliar terminology and complex language can make understanding science difficult for readers.
Approach: They propose a task and dataset for defining scientific terms and controlling the complexity of generated definitions by a sequence-to-sequence approach.
Outcome: The proposed system is based on a sequence-to-sequence approach and human evaluations show it offers superior fluency while controlling complexity.
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation (2024.lrec-main)

Copied to clipboard

Challenge: Increasing misinformation has led to a decrease in trust in news organizations and a decline in the health and medical industry.
Approach: They propose a novel annotation scheme that incorporates persuasive writing tactics in textual documents to aid the automatic identification of misinformation.
Outcome: The proposed scheme improves accuracy and explainability of misinformation detection models.
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context.
Approach: They propose a task of context-aware text generation in the scientific domain to exploit the contributions of context in generated texts.
Outcome: The proposed dataset comprehensively benchmarks the efficacy of the proposed dataset in generating description and paragraph.
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights (2024.findings-emnlp)

Copied to clipboard

Challenge: despite its importance, little work exists on modelling rigour in scientific writing . despite widespread use of term, scientific literature lacks definition of rigor .
Approach: They propose a framework to automatically identify and define rigour criteria and assess their relevance in scientific writing.
Outcome: The proposed framework can be tailored to the evaluation of scientific rigour for different areas.
Presentation Matters: How to Communicate Science in the NLP Venues and in the Wild? (2024.acl-tutorials)

Copied to clipboard

Challenge: a tutorial on communication skills is being proposed to help early career researchers . the tutorial would cover writing, oral presentation and social media presence .
Approach: a tutorial on communication skills is proposed to help early career researchers . the tutorial would cover writing, oral presentation and social media presence .
Outcome: a new tutorial will cover communication skills, including writing, oral presentation and social media presence . the tutorial will allow attendees to ask questions and clarify their research .
Will This Idea Spread Beyond Academia? Understanding Knowledge Transfer of Scientific Concepts across Text Corpora (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing research on knowledge transfer focuses on documents as unit of analysis and follow their transfer into practice for a specific scientific domain.
Approach: They analyze scientific concepts from corpora and use them to predict knowledge transfer . they find that only a small proportion of these ideas will be used in inventions .
Outcome: The proposed model predicts the use of scientific concepts in clinical trials and inventions.
Towards a Human-Computer Collaborative Scientific Paper Lifecycle: A Pilot Study and Hands-On Tutorial (2024.lrec-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an overview of the scientific paper lifecycle . large language models (LLMs) have increasingly played an important role in academic writing .
Approach: They propose to provide an overview of the scientific paper lifecycle using large language models.
Outcome: The tutorial will provide an overview of the scientific paper lifecycle, including scientific literature understanding, experiment development, manuscript draft writing, and finally draft evaluation.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations