A Framework for Fine-Grained Complexity Control in Health Answer Generation (2025.acl-srw)
Copied to clipboard
| Challenge: | Health literacy is the ability to obtain, process, and understand basic health information. |
| Approach: | They propose a framework for automatically generating health answers at multiple, precisely controlled complexity levels. |
| Outcome: | The proposed framework allows users to generate health questions at multiple complexity levels. |
Similar Papers
DocLens: Multi-aspect Fine-grained Medical Text Evaluation (2024.acl-long)
Copied to clipboard
Yiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn Rose
| Challenge: | Medical text generation systems are widely used to assist with administrative work and highlight salient information to support decision-making. |
| Approach: | They propose a set of metrics to evaluate completeness, conciseness, and attribution of medical text at a fine-grained level. |
| Outcome: | The proposed framework exhibits substantially higher agreement with medical experts than existing metrics. |
Simple or Complex? Complexity-controllable Question Generation with Soft Templates and Deep Mixture of Experts Model (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on complex questions does not consider controlling complexity of generated questions. |
| Approach: | They propose an end-to-end neural complexity-controllable question generation model that incorporates a mixture of experts as the selector of soft templates to capture question similarity while avoiding the expensive construction of actual templates. |
| Outcome: | The proposed model is superior to state-of-the-art methods in both automatic and manual evaluations on two benchmark QA datasets. |
Can Large Language Models Accurately Generate Answer Keys for Health-related Questions? (2025.acl-short)
Copied to clipboard
| Challenge: | Evaluating the factuality of LLM generated answers is challenging for many tasks, including question answering. |
| Approach: | They propose to use information nuggets to evaluate the factuality of LLM generated answers . they find providing an example and extracting nuggots from an answer is the best approach . |
| Outcome: | The proposed model performs best when compared to human nugget generation. |
imapScore: Medical Fact Evaluation Made Easy (2024.findings-acl)
Copied to clipboard
| Challenge: | Automated evaluation of natural language generation tasks fails to focus on medical QA because of the diversity in medical terminology. |
| Approach: | They propose a new data structure, imap, to capture key information in questions and answers. |
| Outcome: | The proposed model outperforms state-of-the-art metrics in correlation with human scores. |
MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using fine-grained readability measures is the first step towards making medical texts more accessible. |
| Approach: | They propose a dataset MedReadMe which measures sentences and complex spans with an annotation tool. |
| Outcome: | The proposed dataset covers 650 linguistic features and additional complex span features, and is compared against state-of-the-art methods using large language models. |
Multilingual Simplification of Medical Texts (2023.emnlp-main)
Copied to clipboard
Sebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh Ramanathan, Wei Xu, Byron Wallace, Junyi Jessy Li
| Challenge: | Existing work on medical text simplification has focused on monolingual settings . important findings in medicine are typically presented in technical, jargon-laden language . text simulating models can generate viable simplified texts, but there are outstanding challenges . |
| Approach: | They propose a dataset for medical text simplification in four languages . they evaluate fine-tuned and zero-shot models across these languages based on human assessments and analyses . |
| Outcome: | The proposed dataset evaluates models in English, Spanish, French, and Farsi . it shows that the models can generate viable simplified texts, but there are challenges . |
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions (2025.naacl-long)
Copied to clipboard
| Challenge: | Medical board exams or general clinical questions do not capture the complexity of real clinical cases. |
| Approach: | They construct two datasets that are structured as multiple-choice question-answering tasks accompanied by expert-written explanations. |
| Outcome: | The proposed datasets are harder than previous benchmarks. |
Trustworthy Medical Question Answering: An Evaluation-Centric Survey (2025.emnlp-main)
Copied to clipboard
Yinuo Wang, Baiyang Wang, Robert Mercer, Frank Rudzicz, Sudipta Singha Roy, Pengjie Ren, Zhumin Chen, Xindi Wang
| Challenge: | achieving comprehensive trustworthiness in medical QA poses significant challenges due to complexity of healthcare data, critical nature of clinical scenarios, and multifaceted dimensions of trustworthy AI. |
| Approach: | They examine six key dimensions of trustworthiness in medical QA . they compare how each dimension is evaluated in existing LLM-based systems . |
| Outcome: | The findings show that large language models have improved patient safety and effectiveness . the models exhibit critical trust failures when deployed in clinical settings . |
Controlling Pre-trained Language Models for Grade-Specific Text Simplification (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to text simplification control output complexity at corpus level disregarding complexity of individual inputs and considering only one level of output complexity. |
| Approach: | They propose a method that predicts edit operations required for a specific grade level . they say this approach improves the quality of the simplified outputs over corpus-level heuristics . |
| Outcome: | The proposed method improves the readability of simplified outputs over corpus-level search-based heuristics. |
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Infodemics and health misinformation have significant negative impact on individuals and society . generative AI has significantly accelerated the spread and expanded the reach of health misinfo . |
| Approach: | MM-Health is a large scale multimodal misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated multiplemodal information . |
| Outcome: | MM-Health is a large scale misinformation dataset in the health domain . it includes human-generated multimodal information and AI-generated content . |