Challenge: Existing methods for aligning open-ended outputs with fine-grained clinician preferences are weakly grounded in professional guidelines.
Approach: They propose a framework to align large language models' outputs with fine-grained clinician preferences . they propose 119 broadly reusable, clinically grounded principles organized by clinical dimensions .
Outcome: The proposed framework outperforms existing models on HealthBench-Hard and Deepseek-R1 and o3.

Similar Papers

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are difficult to align with high-stakes medical standards due to dissonance between coarse-grained preference signals and complex protocols.
Approach: They propose a framework that aligns Large Language Models with medical standards . they use a dataset generated via a human-in-the-loop pipeline to augment medical instructions .
Outcome: The proposed framework disentangles safety constraints from general proficiency, enabling precise guidance during reinforcement learning.
Toward Reliable Clinical Coding with Language Models: Verification and Lightweight Adaptation (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for clinical code verification fail to account for hierarchical misalignments . standardized coding systems such as ICD-10-CM1 ensure consistency across medical records.
Approach: They propose to use prompt engineering and small-scale fine-tuning to improve accuracy without the computational overhead of search-based methods.
Outcome: The proposed task is a standalone task and a pipeline component to address hierarchical near-miss errors without the computational overhead of search-based methods.
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment (2025.acl-long)

Copied to clipboard

Challenge: Large language models such as GPT-4 have limited their deployment in clinical settings . a novel framework for adapting SLMs into high-performing clinical models is needed .
Approach: They propose a framework for adapting large language models into high-performing clinical models . they pre-instruct experts on relevant medical and clinical corpora and model merging .
Outcome: The proposed framework outperforms the existing model on the CLUE+ benchmark on medical entities and radiology reports.
A Grounded Preference Model for LLM Alignment (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from factual inconsistency and hallucination despite recent advances . training a preference model requires substantial human annotation, which is expensive and labor-intensive.
Approach: They propose to generate synthetic grounded preference data and train a Grounded Preference Model to assess the overall quality of grounded responses.
Outcome: The proposed model can generate much better grounded responses as judged by GPT4 and achieves the TRUE faithfulness Benchmark.
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
Pedagogical Alignment of Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often used without pedagogical fine-tuning and provide immediate answers rather than guiding students through the problem-solving process.
Approach: They propose a method for constructing large-scale preference datasets using synthetic data generation techniques that eliminates the need for manual annotation.
Outcome: The proposed methods outperform standard supervised fine-tuning (SFT) and improve alignment accuracy by 13.1% and 8.7% respectively.
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios (2024.emnlp-main)

Copied to clipboard

Challenge: Chinese medical large language models (LLMs) are underperforming on this benchmark, especially where medical reasoning and factual consistency are vital.
Approach: They propose a benchmark with 14 expert-guided clinical scenarios to assess the medical ability of large language models across 7 pivot dimensions.
Outcome: The proposed benchmark has been validated in several ways.
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap.
Approach: They propose a benchmark that links MIMIC-IV EHRs to a unified knowledge base built from UMLS and other biomedical vocabularies.
Outcome: The proposed model improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors.
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback (2026.acl-long)

Copied to clipboard

Challenge: Traditional alignment methods rely on human annotations and are subjective and misalignment with real-world user preferences.
Approach: They propose a framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically.
Outcome: The proposed framework identifies and classifies user feedback to LLM responses between conversation turns and creates examples of preferred and dispreferred responses according to user preferences.
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored.
Approach: They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies.
Outcome: The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations