Papers by Bongwon Suh
Blinded by Context: Unveiling the Halo Effect of MLLM in AI Hiring (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Multimodal Large Language Modells (MLLMs) are increasingly being deployed across a range of domains, including finance, law, peer review, and recruitment. |
| Approach: | They investigated how image-based evaluations are influenced by non-job-related information, including extracurricular activities and social media images. |
| Outcome: | The proposed models exhibit significant halo effects in image-based evaluations while text-based assessments showed more resistance to bias. |
Will LLMs Sink or Swim? Exploring Decision-Making Under Pressure (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have shown their ability to simulate human-like decision-making, yet the impact of psychological pressures on their decision- making processes remains underexplored. |
| Approach: | They used explicit and implicit pressure prompts to induce specific pressures and tested them on reasoning, psychometric, and game theory tasks. |
| Outcome: | The results show that pressures significantly affect LLMs’ decision-making, varying across tasks and models. |
Whose Voice, Whose Avatar? Gender Matching Bias in Multimodal AI Teammates (2026.findings-acl)
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly deployed as social agents . yet their ability to integrate conflicting identity cues remains underexplored . |
| Approach: | They audit gender bias in MLLMs that pair synthetic voices with avatars of varying gender presentation and visual fidelity. |
| Outcome: | The findings show that multimodal fairness is not monolithic . they show that models may appear unbiased on one dimension while enforcing stereotypes on another . |
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing function-calling benchmarks focus on single-turn interactions but ignore complexity of real-world scenarios. |
| Approach: | They propose a framework that constructs practical function-calling datasets by synthesizing conversations through a tool graph that maintains dependencies across rounds. |
| Outcome: | The proposed framework synthesizes conversations through a tool graph that maintains dependencies across rounds and a multi-agent system with distinct personas to enhance dialogue naturalness. |
Visual Interference in Speech Evaluation: Cultural Asymmetry and Cross-Modal Bias in MLLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | a new paradigm shifts the paradigm of speech processing from simple transcription to complex social reasoning. |
| Approach: | They construct a cross-modal dataset to examine cultural asymmetry in MLLMs . they find that ML models actively reproduce context-dependent sociolinguistic ideologies based on native audio . |
| Outcome: | The proposed model exhibits cultural asymmetry in anglophone and Korean contexts . the model reproduces sociolinguistic ideologies, consistent with Expectancy Violation Theory . |
Feeling Right vs. Being Right: How AI Sycophancy Affects Value-Laden Deliberation (2026.acl-long)
Copied to clipboard
| Challenge: | Unlike human flattery, AI sycophancy is intentional and self-interested . scophancies are a byproduct of RLHF's user-preference alignment process . |
| Approach: | They propose to operationalize AI sycophancy as excessive face-saving, either active (preserving positive face through agreement) or passive (preserving negative face by withholding challenge). |
| Outcome: | The findings show that sycophancy is a byproduct of RLHF's user-preference alignment process and that it is not a human trait. |
CliniCAST: Benchmarking Acoustic Grounding and Text Dominance in Medical Triage (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Audio-Language Models (LALMs) integrate acoustic capabilities into reasoning, yet whether they reliably ground clinical judgments in audible evidence remains unproven. |
| Approach: | They propose a benchmark that disentangles clinically meaningful acoustic cues from lexical content and speaker demographics. |
| Outcome: | Evaluating 5,856 synthetic samples across 12 disease conditions, the proposed model exhibits fragile acoustic grounding and pronounced "text dominance" failure mode. |