Challenge: Fig. 6: Annotation of biomedical abstracts for automatic detection of inadequate claims (spin) spin is a misleading presentation of scientific results in randomized controlled trials, an important type of clinical trial.
Approach: They propose an algorithm for automatic detection of inadequate claims (spin) they propose to use a corpus of biomedical articles for the task .
Outcome: The proposed algorithm can detect inadequate claims in biomedical abstracts without requiring any prior knowledge of the literature.

Similar Papers

Inferring Which Medical Treatments Work from Reports of Clinical Trials (N19-1)

Copied to clipboard

Challenge: Ideally, one would consult all available evidence from relevant clinical trials. however, these results are primarily disseminated in natural language scientific articles.
Approach: They propose a task that involves inferring results from a full-text article describing randomized controlled trials with respect to a given intervention, comparator, and outcome of interest.
Outcome: The proposed task consists of 10,000+ prompts coupled with full-text articles describing randomized controlled trials.
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)

Copied to clipboard

Challenge: In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day.
Approach: They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured.
Outcome: The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials.
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)

Copied to clipboard

Challenge: Only very few annotated corpora in the medical domain exist.
Approach: They propose to annotate medical entities in case reports from PubMed Central's open access library.
Outcome: The proposed corpus is the first of its kind to be made available to the scientific community in English.
Annotation of a Large Clinical Entity Corpus (D18-1)

Copied to clipboard

Challenge: Past researches have shown the superiority of statistical/ML approaches over the rule based approaches.
Approach: They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning.
Outcome: The proposed corpus contains 5,160 clinical documents from forty different clinical specialties.
A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studies (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey analyses 198 studies published between January 2022 and March 2025 .
Approach: This survey synthesizes recent advances in CV corpus creation and system design.
Outcome: The results of this study are synthesized from 198 studies published between January 2022 and March 2025.
Fair Evaluation in Concept Normalization: a Large-scale Comparative Analysis for BERT-based Models (2020.coling-main)

Copied to clipboard

Challenge: a large number of biomedical entity mentions are retrieved from different ontologies, requiring non-syntactic interpretation.
Approach: They propose to use bidirectional encoder representations from transformers to link biomedical entities across three domains for a task called medical concept normalization.
Outcome: The proposed neural architectures are efficient for linking biomedical entities across domains and corpora.
MedCATTrainer: A Biomedical Free Text Annotation Interface with Active Learning and Research Use Case Specific Customisation (D19-3)

Copied to clipboard

Challenge: 80% of biomedical data is stored in unstructured text such as electronic health records (EHRs).
Approach: They propose a web-based interface for building, improving and customising a given Named Entity Recognition and Linking (NER+L) model for biomedical domain text.
Outcome: The proposed interface is designed to build, improve and customise a NER+L model for biomedical domain text and collate accurate research use case specific training data.
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have raised concerns about reliability and trustworthiness of the models.
Approach: They analyze 134 papers and introduce a taxonomy of evidence-based text generation with LLMs.
Outcome: The proposed methods highlight open challenges and outline promising directions for future work.
NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries (2024.emnlp-main)

Copied to clipboard

Challenge: Despite substantial investment, developing new treatments for neurological conditions is a challenging and often unsuccessful endeavour.
Approach: They propose a corpus for named entity recognition that is annotated clinical trial summaries from ClinicalTrials.gov.
Outcome: The proposed corpus is annotated for neurological diseases, therapeutic interventions, and control treatments and achieves a close-to-human performance.
Towards Understanding of Medical Randomized Controlled Trials by Conclusion Generation (D19-62)

Copied to clipboard

Challenge: Using machine learning to interpret large amounts of data can be over-whelming for clinicians.
Approach: They propose to use PubMed 200k RCT sentence classification dataset to generate RCT conclusion generation task.
Outcome: The proposed model improves quality and correctness in generated conclusions compared to baseline model . the proposed model is not suitable for all RCTs, but it could be improved .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations