Papers by Bongwon Suh

7 papers
Blinded by Context: Unveiling the Halo Effect of MLLM in AI Hiring (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) and Multimodal Large Language Modells (MLLMs) are increasingly being deployed across a range of domains, including finance, law, peer review, and recruitment.
Approach: They investigated how image-based evaluations are influenced by non-job-related information, including extracurricular activities and social media images.
Outcome: The proposed models exhibit significant halo effects in image-based evaluations while text-based assessments showed more resistance to bias.
Will LLMs Sink or Swim? Exploring Decision-Making Under Pressure (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have shown their ability to simulate human-like decision-making, yet the impact of psychological pressures on their decision- making processes remains underexplored.
Approach: They used explicit and implicit pressure prompts to induce specific pressures and tested them on reasoning, psychometric, and game theory tasks.
Outcome: The results show that pressures significantly affect LLMs’ decision-making, varying across tasks and models.
Whose Voice, Whose Avatar? Gender Matching Bias in Multimodal AI Teammates (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models are increasingly deployed as social agents . yet their ability to integrate conflicting identity cues remains underexplored .
Approach: They audit gender bias in MLLMs that pair synthetic voices with avatars of varying gender presentation and visual fidelity.
Outcome: The findings show that multimodal fairness is not monolithic . they show that models may appear unbiased on one dimension while enforcing stereotypes on another .
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues (2025.findings-acl)

Copied to clipboard

Challenge: Existing function-calling benchmarks focus on single-turn interactions but ignore complexity of real-world scenarios.
Approach: They propose a framework that constructs practical function-calling datasets by synthesizing conversations through a tool graph that maintains dependencies across rounds.
Outcome: The proposed framework synthesizes conversations through a tool graph that maintains dependencies across rounds and a multi-agent system with distinct personas to enhance dialogue naturalness.
Visual Interference in Speech Evaluation: Cultural Asymmetry and Cross-Modal Bias in MLLMs (2026.findings-acl)

Copied to clipboard

Challenge: a new paradigm shifts the paradigm of speech processing from simple transcription to complex social reasoning.
Approach: They construct a cross-modal dataset to examine cultural asymmetry in MLLMs . they find that ML models actively reproduce context-dependent sociolinguistic ideologies based on native audio .
Outcome: The proposed model exhibits cultural asymmetry in anglophone and Korean contexts . the model reproduces sociolinguistic ideologies, consistent with Expectancy Violation Theory .
Feeling Right vs. Being Right: How AI Sycophancy Affects Value-Laden Deliberation (2026.acl-long)

Copied to clipboard

Challenge: Unlike human flattery, AI sycophancy is intentional and self-interested . scophancies are a byproduct of RLHF's user-preference alignment process .
Approach: They propose to operationalize AI sycophancy as excessive face-saving, either active (preserving positive face through agreement) or passive (preserving negative face by withholding challenge).
Outcome: The findings show that sycophancy is a byproduct of RLHF's user-preference alignment process and that it is not a human trait.
CliniCAST: Benchmarking Acoustic Grounding and Text Dominance in Medical Triage (2026.findings-acl)

Copied to clipboard

Challenge: Recent Large Audio-Language Models (LALMs) integrate acoustic capabilities into reasoning, yet whether they reliably ground clinical judgments in audible evidence remains unproven.
Approach: They propose a benchmark that disentangles clinically meaningful acoustic cues from lexical content and speaker demographics.
Outcome: Evaluating 5,856 synthetic samples across 12 disease conditions, the proposed model exhibits fragile acoustic grounding and pronounced "text dominance" failure mode.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations