Annotating Spin in Biomedical Scientific Publications : the case of Random Controlled Trials (RCTs) (L18-1)
Copied to clipboard
| Challenge: | Fig. 6: Annotation of biomedical abstracts for automatic detection of inadequate claims (spin) spin is a misleading presentation of scientific results in randomized controlled trials, an important type of clinical trial. |
| Approach: | They propose an algorithm for automatic detection of inadequate claims (spin) they propose to use a corpus of biomedical articles for the task . |
| Outcome: | The proposed algorithm can detect inadequate claims in biomedical abstracts without requiring any prior knowledge of the literature. |
Similar Papers
Inferring Which Medical Treatments Work from Reports of Clinical Trials (N19-1)
Copied to clipboard
| Challenge: | Ideally, one would consult all available evidence from relevant clinical trials. however, these results are primarily disseminated in natural language scientific articles. |
| Approach: | They propose a task that involves inferring results from a full-text article describing randomized controlled trials with respect to a given intervention, comparator, and outcome of interest. |
| Outcome: | The proposed task consists of 10,000+ prompts coupled with full-text articles describing randomized controlled trials. |
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)
Copied to clipboard
| Challenge: | In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day. |
| Approach: | They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured. |
| Outcome: | The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials. |
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)
Copied to clipboard
| Challenge: | Only very few annotated corpora in the medical domain exist. |
| Approach: | They propose to annotate medical entities in case reports from PubMed Central's open access library. |
| Outcome: | The proposed corpus is the first of its kind to be made available to the scientific community in English. |
Annotation of a Large Clinical Entity Corpus (D18-1)
Copied to clipboard
| Challenge: | Past researches have shown the superiority of statistical/ML approaches over the rule based approaches. |
| Approach: | They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning. |
| Outcome: | The proposed corpus contains 5,160 clinical documents from forty different clinical specialties. |
A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studies (2025.findings-emnlp)
Copied to clipboard
| Challenge: | This survey analyses 198 studies published between January 2022 and March 2025 . |
| Approach: | This survey synthesizes recent advances in CV corpus creation and system design. |
| Outcome: | The results of this study are synthesized from 198 studies published between January 2022 and March 2025. |
Fair Evaluation in Concept Normalization: a Large-scale Comparative Analysis for BERT-based Models (2020.coling-main)
Copied to clipboard
| Challenge: | a large number of biomedical entity mentions are retrieved from different ontologies, requiring non-syntactic interpretation. |
| Approach: | They propose to use bidirectional encoder representations from transformers to link biomedical entities across three domains for a task called medical concept normalization. |
| Outcome: | The proposed neural architectures are efficient for linking biomedical entities across domains and corpora. |
MedCATTrainer: A Biomedical Free Text Annotation Interface with Active Learning and Research Use Case Specific Customisation (D19-3)
Copied to clipboard
| Challenge: | 80% of biomedical data is stored in unstructured text such as electronic health records (EHRs). |
| Approach: | They propose a web-based interface for building, improving and customising a given Named Entity Recognition and Linking (NER+L) model for biomedical domain text. |
| Outcome: | The proposed interface is designed to build, improve and customise a NER+L model for biomedical domain text and collate accurate research use case specific training data. |
Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have raised concerns about reliability and trustworthiness of the models. |
| Approach: | They analyze 134 papers and introduce a taxonomy of evidence-based text generation with LLMs. |
| Outcome: | The proposed methods highlight open challenges and outline promising directions for future work. |
NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries (2024.emnlp-main)
Copied to clipboard
Simona Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman, Amelia Cannon, Gerold Schneider, Benjamin Ineichen
| Challenge: | Despite substantial investment, developing new treatments for neurological conditions is a challenging and often unsuccessful endeavour. |
| Approach: | They propose a corpus for named entity recognition that is annotated clinical trial summaries from ClinicalTrials.gov. |
| Outcome: | The proposed corpus is annotated for neurological diseases, therapeutic interventions, and control treatments and achieves a close-to-human performance. |
Towards Understanding of Medical Randomized Controlled Trials by Conclusion Generation (D19-62)
Copied to clipboard
| Challenge: | Using machine learning to interpret large amounts of data can be over-whelming for clinicians. |
| Approach: | They propose to use PubMed 200k RCT sentence classification dataset to generate RCT conclusion generation task. |
| Outcome: | The proposed model improves quality and correctness in generated conclusions compared to baseline model . the proposed model is not suitable for all RCTs, but it could be improved . |