Challenge: Existing annotations for other NLP tasks are used to generate domain-specific large-scale question answering (QA) datasets.
Approach: They propose to re-purpose existing annotations for other NLP tasks by generating a large-scale question answering corpus using 1 million questions-logical form and 400,000+ question-answer evidence pairs.
Outcome: The proposed model can be trained to learn domain-specific large-scale question answering (QA) datasets.

Similar Papers

DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries (2022.lrec-1)

Copied to clipboard

Challenge: a new question answering dataset is being developed for electronic health records . structured tables and unstructured notes can be duplicated, contradictory or provide additional context .
Approach: They develop a question-answer-matching dataset using structured tables and unstructured notes from an EHR.
Outcome: The proposed model is based on a model with a modality selection network . it uses the prediction of a RAT-SQL to choose between EHR tables and clinical notes .
Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset (2020.acl-main)

Copied to clipboard

Challenge: Medical professionals often query over clinical notes to find information that can support their decision making.
Approach: They propose to use expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering based on clinical notes.
Outcome: The proposed system can answer clinical questions without using domain knowledge.
Large-Scale QA-SRL Parsing (P18-1)

Copied to clipboard

Challenge: a crowd-sourced approach to learning semantic parsers to predict predicateargument structures is open to many researchers.
Approach: They propose a large-scale corpus of Question-Answer driven Semantic Role Labeling annotations . they also propose QA-SRL Bank 2.0, a crowd-sourcing scheme that can be used to train high quality parsers .
Outcome: The proposed QA-SRL parser can generate high-quality questions at low cost and is intuitive to non-experts.
MedREQAL: Examining Medical Knowledge Recall of Large Language Models via Question Answering (2024.findings-acl)

Copied to clipboard

Challenge: Large language models can encode knowledge during pre-training on large text corpora, enabling downstream tasks like question answering (QA).
Approach: They construct a dataset derived from systematic reviews to examine their ability to encode medical knowledge and their recall.
Outcome: The proposed model performs well on the biomedical QA dataset.
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems (2024.acl-demos)

Copied to clipboard

Challenge: Existing tools for augmented question-answering do not support researchers and developers to customize the training, testing, and deployment process.
Approach: They propose an open-source toolkit that features a wide selection of model training algorithms, evaluation methods, and deployment tools curated from the latest research.
Outcome: The proposed framework trains and deploys 7B-models with the same performance as OpenAI’s text-ada-002 and GPT-4-turbo.
UQA: Corpus for Urdu Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Urdu is a low-resource language with over 70 million native speakers . expanding the reach of NLP to languages other than English is crucial for advancing multilingual AI systems.
Approach: They introduce a novel dataset for question answering and text comprehension in Urdu . they use a technique called EATS which preserves the answer spans in translated context paragraphs .
Outcome: The proposed dataset preserves answer spans in translated context paragraphs.
PAXQA: Generating Cross-lingual Question Answering Examples at Training Scale (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing question answering systems rely on large, high-quality training data.
Approach: They propose a synthetic data generation method which decomposes cross-lingual QA into two stages . they apply a question generation model to the English side and annotation projection to translate both questions and answers.
Outcome: The proposed method outperforms existing methods on cross-lingual QA datasets.
RadQA: A Question Answering Dataset to Improve Comprehension of Radiology Reports (2022.lrec-1)

Copied to clipboard

Challenge: Question answering (QA) is an intuitive means to query text data.
Approach: They propose a radiology question-answer-evidence-pair dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians.
Outcome: The proposed dataset has 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians.
A Framework for Automatic Generation of Spoken Question-Answering Data (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing frameworks to automatically generate a spoken question answering dataset are limited by the amount of spoken text documents available.
Approach: They propose to use QG module to generate questions from text documents, TTS module to convert text documents into spoken form and automatic speech recognition module to transcribe spoken content.
Outcome: The proposed framework is efficient for automatically generating spoken QA datasets.
ELI5: Long Form Question Answering (P19-1)

Copied to clipboard

Challenge: Existing question answering datasets provide extractive or short answers, but less attention has been paid to open-ended questions that require explanations.
Approach: They present a large-scale corpus for long form question answering . they use a Reddit forum to provide elaborate answers to open-ended questions .
Outcome: The proposed model outperforms Seq2Seq, language modeling, and other models in human evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations