Papers by Hyunwoo Yoo
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections (2026.findings-acl)
Copied to clipboard
| Challenge: | Several studies rely on additional models to optimize mixtures. |
| Approach: | They propose a method that dynamically optimizes instruction-tuning dataset mixtures by prior-scaled Boltzmann Exploration and a multi-armed bandit setup. |
| Outcome: | The proposed method improves the TÜLU-2-mixture and TÜLO-3-mixtures across 10 benchmarks while introducing minimal computational overhead over naive sampling. |
Whose Voice, Whose Avatar? Gender Matching Bias in Multimodal AI Teammates (2026.findings-acl)
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly deployed as social agents . yet their ability to integrate conflicting identity cues remains underexplored . |
| Approach: | They audit gender bias in MLLMs that pair synthetic voices with avatars of varying gender presentation and visual fidelity. |
| Outcome: | The findings show that multimodal fairness is not monolithic . they show that models may appear unbiased on one dimension while enforcing stereotypes on another . |
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive performance on downstream tasks, but if they cannot be fully described in prompts, they could fail to perform the task. |
| Approach: | They propose a method to contextualize a task toward a large language model (LLM) they use open-ended zero-shot inference from the entire dataset to aggregate the inference results and incorporate the aggregated meta-information for the actual task. |
| Outcome: | The proposed method improves text clustering tasks and improves on several datasets. |
Visual Interference in Speech Evaluation: Cultural Asymmetry and Cross-Modal Bias in MLLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | a new paradigm shifts the paradigm of speech processing from simple transcription to complex social reasoning. |
| Approach: | They construct a cross-modal dataset to examine cultural asymmetry in MLLMs . they find that ML models actively reproduce context-dependent sociolinguistic ideologies based on native audio . |
| Outcome: | The proposed model exhibits cultural asymmetry in anglophone and Korean contexts . the model reproduces sociolinguistic ideologies, consistent with Expectancy Violation Theory . |
CliniCAST: Benchmarking Acoustic Grounding and Text Dominance in Medical Triage (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Audio-Language Models (LALMs) integrate acoustic capabilities into reasoning, yet whether they reliably ground clinical judgments in audible evidence remains unproven. |
| Approach: | They propose a benchmark that disentangles clinically meaningful acoustic cues from lexical content and speaker demographics. |
| Outcome: | Evaluating 5,856 synthetic samples across 12 disease conditions, the proposed model exhibits fragile acoustic grounding and pronounced "text dominance" failure mode. |