BioRead: A New Dataset for Biomedical Reading Comprehension (L18-1)

Copied to clipboard

Challenge: BioRead is a publicly available cloze-style biomedical machine reading comprehension (MRC) dataset with 16.4 million passage-question instances.
Approach: They propose to build a cloze-style biomedical machine reading comprehension (MRC) dataset with 16.4 million passage-question instances.
Outcome: The proposed method outperforms baselines on bioReadLite and bioASQ, and is currently the best on BioReadLite.

Similar Papers

ScholarlyRead: A New Dataset for Scientific Article Reading Comprehension (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on MRC on scholarly articles have focused on general domain datasets of news articles and elementary school-level storybooks.
Approach: They propose to generate automatic questions from span-of-word-based scholarly articles’ Reading Comprehension dataset with approximately 10K manually checked passage-question-answer instances.
Outcome: The proposed model yields the F1 score of 37.31% and is useful for building Question-Answering (QA) systems on scientific articles.
CliCR: a Dataset of Clinical Case Reports for Machine Reading Comprehension (N18-1)

Copied to clipboard

Challenge: Currently, machine comprehension datasets are extremely scarce for specialized domains.
Approach: They propose a dataset for machine comprehension in the medical domain using clinical case reports with around 100,000 gap-filling queries about these cases.
Outcome: The proposed dataset uses clinical case reports with around 100,000 gap-filling queries about these cases.
Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset (2020.acl-main)

Copied to clipboard

Challenge: Medical professionals often query over clinical notes to find information that can support their decision making.
Approach: They propose to use expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering based on clinical notes.
Outcome: The proposed system can answer clinical questions without using domain knowledge.
Dataset for the First Evaluation on Chinese Machine Reading Comprehension (L18-1)

Copied to clipboard

Challenge: Existing reading comprehension datasets are mostly in English .
Approach: They propose a Chinese reading comprehension dataset to add diversity to existing reading comprehension data . proposed dataset contains cloze-style reading comprehension and user query reading comprehension .
Outcome: The proposed dataset is based on a Chinese reading comprehension dataset . it includes two types of cloze-style and user query reading comprehension . the proposed dataset hosted the 1st Evaluation on Chinese Machine Reading Comprehension (CMRC-2017)
Towards Medical Machine Reading Comprehension with Structural Knowledge and Plain Text (2020.emnlp-main)

Copied to clipboard

Challenge: MRC has achieved significant progress on the open domain in recent years due to large-scale pre-trained language models.
Approach: They propose a machine reading comprehension model which exploits structural medical knowledge and reference medical plain text to improve the exam's accuracy.
Outcome: The proposed model outperforms existing models with a large margin and passes the exam with 61.8% accuracy rate on the test set.
BioReader: a Retrieval-Enhanced Text-to-Text Transformer for Biomedical Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Recent research has equipped language models with the ability to attend over relevant and factual information from non-parametric external sources, drawing a complementary path to architectural scaling.
Approach: They propose a retrieval-enhanced text-to-text model that augments the input prompt by fetching and assembling relevant scientific literature chunks from a neural database centered on PubMed.
Outcome: The proposed model outperforms state-of-the-art models on a broad array of downstream tasks while using up to 3x fewer parameters.
Comprehensive Multi-Dataset Evaluation of Reading Comprehension (D19-58)

Copied to clipboard

Challenge: Recent research aims to facilitate training and evaluation on several reading comprehension datasets at the same time.
Approach: They propose an evaluation server that reports performance on seven diverse reading comprehension datasets and includes synthetic augmentations to test models' ability to handle out-of-domain questions.
Outcome: The evaluation server performs on seven reading comprehension datasets, and collects and includes synthetic augmentations for these datasets to test models' ability to handle out-of-domain questions.
English Machine Reading Comprehension Datasets: A Survey (2021.emnlp-main)

Copied to clipboard

Challenge: a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape .
Approach: They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem.
Outcome: The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem.
M-QALM: A Benchmark to Assess Clinical Reading Comprehension and Knowledge Recall in Large Language Models via Question Answering (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on adapting large language models to perform a variety of tasks in high-stakes domains such as healthcare lack understanding of the extent and contributing factors that allow them to recall relevant knowledge and combine it with presented information.
Approach: They propose to use multiple choice and abstractive question answering to investigate the extent and contributing factors that allow LLMs to recall relevant knowledge and combine it with presented information in the clinical and biomedical domain.
Outcome: The proposed models perform better on 22 datasets in three generalist and three specialist biomedical sub-domains, and show that they can generalise to unseen sub- domains.
Improving Machine Reading Comprehension with General Reading Strategies (N19-1)

Copied to clipboard

Challenge: Recent studies have shown that reading strategies improve comprehension levels for readers lacking adequate prior knowledge.
Approach: They propose three general strategies to improve machine reading comprehension (MRC) by fine-tuning a pre-trained model with strategies and a target task.
Outcome: The proposed models improve non-extractive machine reading comprehension (MRC) on the largest general domain multiple-choice dataset RACE.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations