| Challenge: | In this paper, we establish a PICO and a sentiment annotated corpus of clinical trial publications. |
| Approach: | They propose to create a phrase-level PICO corpus and a sentence-level sentiment annotated corpus from clinical trial publications. |
| Outcome: | The proposed corpus is annotated on a phrase-level and a sentiment annotation on the same corpus. |
Similar Papers
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)
Copied to clipboard
| Challenge: | In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day. |
| Approach: | They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured. |
| Outcome: | The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials. |
Named Entities in Medical Case Reports: Corpus and Experiments (2020.lrec-1)
Copied to clipboard
| Challenge: | Only very few annotated corpora in the medical domain exist. |
| Approach: | They propose to annotate medical entities in case reports from PubMed Central's open access library. |
| Outcome: | The proposed corpus is the first of its kind to be made available to the scientific community in English. |
Annotation of a Large Clinical Entity Corpus (D18-1)
Copied to clipboard
| Challenge: | Past researches have shown the superiority of statistical/ML approaches over the rule based approaches. |
| Approach: | They propose to annotate a clinical domain annotated corpus using a small data set or a narrower domain to take full advantage of machine learning. |
| Outcome: | The proposed corpus contains 5,160 clinical documents from forty different clinical specialties. |
The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)
Copied to clipboard
Izhak Shafran, Nan Du, Linh Tran, Amanda Perry, Lauren Keyes, Mark Knichel, Ashley Domin, Lei Huang, Yu-hui Chen, Gang Li, Mingqiu Wang, Laurent El Shafey, Hagen Soltau, Justin Stuart Paul
| Challenge: | Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking. |
| Approach: | They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model. |
| Outcome: | The proposed model is more useful than the F-scores reflect and can be used in clinical notes. |
An Annotated Corpus of Textual Explanations for Clinical Decision Support (2022.lrec-1)
Copied to clipboard
Roland Roller, Aljoscha Burchardt, Nils Feldhus, Laura Seiffe, Klemens Budde, Simon Ronicke, Bilgin Osmanodja
| Challenge: | In recent years, machine learning for clinical decision support has gained more and more attention. |
| Approach: | They propose to use XAI to provide an explanation of a model's decision making process by constructing a corpus of sentences that are annotated with different semantic layers. |
| Outcome: | The proposed models outperform physicians on very specific, narrow tasks or can help physicians to work more efficiently. |
COMETA: A Corpus for Medical Entity Linking in the Social Media (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for Entity Linking (EL) fail to address the complex nature of health terminology in layman’s language. |
| Approach: | They propose to use a corpus of 20k English biomedical entity mentions from Reddit expert-annotated with links to a widely-used medical knowledge graph to investigate the ability of these systems to perform complex inference on entities and concepts. |
| Outcome: | The proposed corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by existing resources in the field. |
Sent2Span: Span Detection for PICO Extraction in the Biomedical Text without Span Annotations (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Experiments show that PICO span detection results achieve much higher results for recall when compared to fully supervised methods. |
| Approach: | They propose to extract and then normalise PICO information from clinical trial articles and use crowdsourced sentence-level annotations to detect spans. |
| Outcome: | The proposed method achieves much higher results for recall when compared to fully supervised methods with PICO sentence detection at least as good as human annotations. |
Medical Crossing: a Cross-lingual Evaluation of Clinical Entity Linking (2022.lrec-1)
Copied to clipboard
Anton Alekseev, Zulfat Miftahutdinov, Elena Tutubalina, Artem Shelmanov, Vladimir Ivanov, Vladimir Kokh, Alexander Nesterov, Manvel Avetisian, Andrei Chertok, Sergey Nikolenko
| Challenge: | Existing approaches to medical entity linking are limited in terms of data volume and languages. |
| Approach: | They propose to use clinical reports, clinical guidelines, and medical research papers to evaluate cross-lingual medical entity linking. |
| Outcome: | The proposed model outperforms existing models on clinical reports, clinical guidelines, and medical research papers. |
A FrameNet for Cancer Information in Clinical Narratives: Schema and Annotation (L18-1)
Copied to clipboard
| Challenge: | Existing natural language processing (NLP) systems for cancer-related information are highly task-specific and often produce incompatible annotations and algorithms. |
| Approach: | They propose a general-purpose natural language processing resource for cancer-related information in clinical notes . the project uses a frame semantic method to emphasize the information presented in the notes themselves . |
| Outcome: | The proposed project emphasizes the information presented in the clinical notes and its linguistic structure. |
An annotated dataset of literary entities (N19-1)
Copied to clipboard
| Challenge: | Existing datasets built on news focus on non-named entities, but not literary texts. |
| Approach: | They propose to annotate 210,532 tokens from 100 different English-language literary texts for ACE entity categories (person, location, geo-political entity, facility, organization, and vehicle). |
| Outcome: | The proposed dataset includes 210,532 tokens drawn from 100 different English-language literary texts. |