Challenge: Health literacy is the ability to obtain, process, and understand basic health information.
Approach: They propose a framework for automatically generating health answers at multiple, precisely controlled complexity levels.
Outcome: The proposed framework allows users to generate health questions at multiple complexity levels.

Similar Papers

DocLens: Multi-aspect Fine-grained Medical Text Evaluation (2024.acl-long)

Copied to clipboard

Challenge: Medical text generation systems are widely used to assist with administrative work and highlight salient information to support decision-making.
Approach: They propose a set of metrics to evaluate completeness, conciseness, and attribution of medical text at a fine-grained level.
Outcome: The proposed framework exhibits substantially higher agreement with medical experts than existing metrics.
Simple or Complex? Complexity-controllable Question Generation with Soft Templates and Deep Mixture of Experts Model (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing work on complex questions does not consider controlling complexity of generated questions.
Approach: They propose an end-to-end neural complexity-controllable question generation model that incorporates a mixture of experts as the selector of soft templates to capture question similarity while avoiding the expensive construction of actual templates.
Outcome: The proposed model is superior to state-of-the-art methods in both automatic and manual evaluations on two benchmark QA datasets.
Can Large Language Models Accurately Generate Answer Keys for Health-related Questions? (2025.acl-short)

Copied to clipboard

Challenge: Evaluating the factuality of LLM generated answers is challenging for many tasks, including question answering.
Approach: They propose to use information nuggets to evaluate the factuality of LLM generated answers . they find providing an example and extracting nuggots from an answer is the best approach .
Outcome: The proposed model performs best when compared to human nugget generation.
imapScore: Medical Fact Evaluation Made Easy (2024.findings-acl)

Copied to clipboard

Challenge: Automated evaluation of natural language generation tasks fails to focus on medical QA because of the diversity in medical terminology.
Approach: They propose a new data structure, imap, to capture key information in questions and answers.
Outcome: The proposed model outperforms state-of-the-art metrics in correlation with human scores.
MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain (2024.emnlp-main)

Copied to clipboard

Challenge: Using fine-grained readability measures is the first step towards making medical texts more accessible.
Approach: They propose a dataset MedReadMe which measures sentences and complex spans with an annotation tool.
Outcome: The proposed dataset covers 650 linguistic features and additional complex span features, and is compared against state-of-the-art methods using large language models.
Multilingual Simplification of Medical Texts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work on medical text simplification has focused on monolingual settings . important findings in medicine are typically presented in technical, jargon-laden language . text simulating models can generate viable simplified texts, but there are outstanding challenges .
Approach: They propose a dataset for medical text simplification in four languages . they evaluate fine-tuned and zero-shot models across these languages based on human assessments and analyses .
Outcome: The proposed dataset evaluates models in English, Spanish, French, and Farsi . it shows that the models can generate viable simplified texts, but there are challenges .
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions (2025.naacl-long)

Copied to clipboard

Challenge: Medical board exams or general clinical questions do not capture the complexity of real clinical cases.
Approach: They construct two datasets that are structured as multiple-choice question-answering tasks accompanied by expert-written explanations.
Outcome: The proposed datasets are harder than previous benchmarks.
Trustworthy Medical Question Answering: An Evaluation-Centric Survey (2025.emnlp-main)

Copied to clipboard

Challenge: achieving comprehensive trustworthiness in medical QA poses significant challenges due to complexity of healthcare data, critical nature of clinical scenarios, and multifaceted dimensions of trustworthy AI.
Approach: They examine six key dimensions of trustworthiness in medical QA . they compare how each dimension is evaluated in existing LLM-based systems .
Outcome: The findings show that large language models have improved patient safety and effectiveness . the models exhibit critical trust failures when deployed in clinical settings .
Controlling Pre-trained Language Models for Grade-Specific Text Simplification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to text simplification control output complexity at corpus level disregarding complexity of individual inputs and considering only one level of output complexity.
Approach: They propose a method that predicts edit operations required for a specific grade level . they say this approach improves the quality of the simplified outputs over corpus-level heuristics .
Outcome: The proposed method improves the readability of simplified outputs over corpus-level search-based heuristics.
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation (2025.findings-emnlp)

Copied to clipboard

Challenge: Infodemics and health misinformation have significant negative impact on individuals and society . generative AI has significantly accelerated the spread and expanded the reach of health misinfo .
Approach: MM-Health is a large scale multimodal misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated multiplemodal information .
Outcome: MM-Health is a large scale misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated content .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations