Challenge: Existing SBDH datasets lack detailed annotations and are limited in their availability and coverage.
Approach: They propose a synthetic SBDH annotation dataset with detailed SBDH status, temporal information, and rationale across 15 categories.
Outcome: The proposed dataset outperforms models with no Synth-SBDH training on three tasks using real-world clinical datasets from two distinct hospital settings.

Similar Papers

SDOH-NLI: a Dataset for Inferring Social Determinants of Health from Clinical Notes (2023.findings-emnlp)

Copied to clipboard

Challenge: Social and behavioral determinants of health (SDOH) play a significant role in shaping health outcomes, and extracting these determinant from clinical notes is a first step to help healthcare providers systematically identify opportunities to provide appropriate care and address disparities.
Approach: They propose a dataset that extracts social and behavioral determinants from clinical notes and uses them to form a natural language inference task.
Outcome: The proposed dataset is based on publicly available notes and is more challenging than standard NLI benchmarks.
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods (2024.lrec-main)

Copied to clipboard

Challenge: Social determinants of health (SDoH) are often studied in the electronic health record (EHR) however, there are difficulties in documenting SDoH in a tabular format due to the lack of a comprehensive SDoh tool.
Approach: They propose to annotate social history sections from 1,260 clinical notes from pediatric patients within the University of Washington (UW) hospital system.
Outcome: The proposed corpus captures ten distinct health determinants including living and economic stability, prior trauma, education access, substance use history, and mental health with an overall annotator agreement of 81.9 F1.
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains (2025.emnlp-demos)

Copied to clipboard

Challenge: SynthTextEval is a toolkit for conducting comprehensive evaluations of synthetic text.
Approach: They propose a toolkit for conducting comprehensive evaluations of synthetic text using large language models.
Outcome: The proposed toolkit can be run over any dataset, but it is aimed at two high-stakes domains: healthcare and law.
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models (2025.acl-short)

Copied to clipboard

Challenge: Large language models (LLMs) rely on superficial cues leading to spurious predictions . recent work has highlighted how LLMs exploit spurious patterns rather than learning causal, generalizable features.
Approach: They use a social history annotation corpus dataset to examine drug status extraction . they evaluate prompt engineering and chain-of-thought reasoning to reduce false positives .
Outcome: The proposed model can predict drug use when alcohol or smoking is not present, while uncovering gender disparities in model performance.
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs (2026.findings-acl)

Copied to clipboard

Challenge: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly.
Approach: They define the first HSS domain system covering 14 mainstream fields and introduce HSS-Synth.
Outcome: the proposed pipeline outperforms 14 leading baselines on 16 benchmarks.
A Distant Supervision Corpus for Extracting Biomedical Relationships Between Chemicals, Diseases and Genes (2022.lrec-1)

Copied to clipboard

Challenge: Biomedical researchers have used manual curation to extract biomedical interactions from research texts to improve coverage.
Approach: They propose a new dataset for training and evaluating multi-class multi-label biomedical relation extraction models using human annotations and the CTD database.
Outcome: The proposed dataset is substantially larger and cleaner than existing datasets and includes annotations linking mentions to their entities.
Can Synthetic Text Help Clinical Named Entity Recognition? A Study of Electronic Health Records in French (2023.eacl-main)

Copied to clipboard

Challenge: In sensitive domains, the sharing of corpora is restricted due to confidentiality, copyrights or trade secrets.
Approach: They use auto-regressive neural models to generate a clinical case corpus annotated with clinical entities and evaluate it for a named entity recognition task.
Outcome: The proposed model can produce clinical case corpus annotated with clinical entities while maintaining confidentiality.
Data-Constrained Synthesis of Training Data for De-Identification (2025.acl-long)

Copied to clipboard

Challenge: sensitive domains lack widely available datasets due to privacy risks . recent studies have focused on evaluating the privacy of the synthetic text .
Approach: They domain-adapt LLMs to clinical domain and generate synthetic clinical texts . they then generate NER models that can be annotated with tags for PII .
Outcome: The proposed model performs better than the original model using smaller datasets.
PersonalityDBench: A Dataset for Personality Disorders - from Modeling to Controlled Generation (2026.acl-long)

Copied to clipboard

Challenge: Personality disorders are chronic, rigid patterns of thinking, behavior, and emotions that deviate from cultural norms and persist in social settings.
Approach: They propose a large-scale, clinically grounded dataset that supports multidimensional study of personality pathology and standardized evaluation of LLM steering toward clinically ground behavioral targets.
Outcome: The PersonalityDBench dataset supports multidimensional study of personality pathology and evaluation of LLM steering toward clinically grounded behavioral targets.
The Medical Scribe: Corpus Development and Model Performance Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to assist in clinical note generation using audio of provider-patient encounters are lacking.
Approach: They develop an annotation scheme to extract relevant clinical concepts from audio of provider-patient encounters and train a state-of-the-art tagging model.
Outcome: The proposed model is more useful than the F-scores reflect and can be used in clinical notes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations