Challenge: Existing medical QA datasets are mostly English and centred on scientific articles or clinical notes.
Approach: They propose an extractive QA dataset built from Riassunti delle Caratteristiche del Prodotto . the final dataset contains 861 high-quality question–answer pairs .
Outcome: The proposed dataset contains 861 high-quality question–answer pairs on indications, contraindications, dosage, warnings, interactions, and pharmacological properties.

Similar Papers

DanteLLM: Let’s Push Italian LLM Research Forward! (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for large language processing in the English language are limited in resources and evaluation tools for non-English languages.
Approach: They propose a benchmark and an open LLM Leaderboard to evaluate LLMs’ performance in Italian and propose 'DanteLLM' it is the most performant LLM in the world, with improvements of up to 6 points .
Outcome: The proposed model outperforms existing models in Italian and offers a blueprint for the development and evaluation of LLMs in other languages.
ITALIC: An Italian Culture-Aware Natural Language Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: ITALIC is a large-scale benchmark dataset of 10,000 multiple-choice questions designed to evaluate the natural language understanding of the Italian language and culture.
Approach: They propose to use a large-scale benchmark dataset to evaluate the natural language understanding of the Italian language and culture.
Outcome: The ITALIC dataset spans 12 domains and uses 17 state-of-the-art LLMs to assess the natural language understanding of the italian language and culture.
ViMedAQA: A Vietnamese Medical Abstractive Question-Answering Dataset and Findings of Large Language Model (2024.acl-srw)

Copied to clipboard

Challenge: Existing abstractive question-answering datasets in Vietnamese are lacking .
Approach: They propose to introduce a Vietnamese abstractive question-answering corpus to address this gap . they propose to use Vietnamese abstractives to generate answers to questions .
Outcome: The proposed dataset examines the capability of large language models in the Vietnamese medical domain, including reasoning, memorizing and awareness of essential information.
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing multiple-choice question and answer (QA) datasets are text-only and available in a limited subset of languages and countries.
Approach: They propose a multilingual, multimodal benchmarking dataset to evaluate multimodal/vision language models in healthcare.
Outcome: The WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries.
Should I Believe in What Medical AI Says? A Chinese Benchmark for Medication Based on Knowledge and Reasoning (2025.acl-short)

Copied to clipboard

Challenge: Large language models (LLMs) generate hallucinations when handling unfamiliar information.
Approach: They propose a Chinese benchmark to evaluate large language models' knowledge and reasoning capabilities in medication tasks.
Outcome: The proposed benchmark evaluates models in indication, dosage and administration, contraindicated population, mechanisms of action, drug recommendation, and drug interaction across six datasets.
Par-ITA: Benchmarking Seq2Seq and LLMs on a Human-Supervised Parallel Corpus for Italian Hyperpartisan Neutralization (2026.acl-long)

Copied to clipboard

Challenge: a new study examines the role of hyperpartisan content in online polarization in the social web.
Approach: They propose a human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs.
Outcome: The proposed dataset is the first human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs.
Truth, Trust, and Trouble: Medical AI on the Edge (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are promising for transforming digital health applications . but ensuring they meet industry standards for factual accuracy, usefulness, and safety remains a challenge .
Approach: They present a framework to assess large language models' accuracy, usefulness, and safety . they assess models' honesty, helpfulness, harmlessness and domain-specific tuning .
Outcome: The proposed framework assesses models across honesty, helpfulness, and harmlessness . AlpaCare-13B achieves highest accuracy (91.7%) and harmlessity (0.92) .
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation (2025.acl-industry)

Copied to clipboard

Challenge: Clinical note generation (CNG) tools are being developed to address extended working hours and healthcare provider fatigue.
Approach: They evaluate the reliability of 12 open-weight and proprietary LLMs from Anthropic, Meta, Mistral, and OpenAI in CNG in terms of their ability to generate notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity)
Outcome: The results show that the LLMs generated notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (symbol similarity) overall, Meta’s Llama 70B was the most reliable, followed by Mistral’s Small model.
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) performance on medical multiplechoice question (MCQ) benchmarks have stimulated interest from healthcare providers and patients globally.
Approach: They introduce AfriMed-QA, the first largescale Pan-African English multi-specialty medical Question-Answering (QA) dataset, with 15,000 questions sourced from over 60 medical schools across 16 countries.
Outcome: The proposed model outperforms other models in the medical field and is compared with other models.
What Does Infect Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs (2026.eacl-long)

Copied to clipboard

Challenge: S-MedQA is an English question-answering dataset designed for benchmarking large language models in fine-grained clinical specialties.
Approach: They propose to use an English medical question-answering dataset to benchmark large language models in clinical specialties.
Outcome: The proposed dataset is designed to benchmark large language models in medical specialties.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations