Papers by Youngjae Kim

28 papers
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild (2025.findings-naacl)

Copied to clipboard

Challenge: EgoSpeak predicts when an agent should begin speaking based on egocentric streaming video.
Approach: They propose a framework for real-time speech initiation prediction in egocentric streaming video by modeling the conversation from the camera wearer's first-person perspective.
Outcome: The proposed framework outperforms random and silence-based baselines in real time and highlights the importance of multimodal input and context length in effectively deciding when to speak.
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety evaluations rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail.
Approach: They propose a framework for systematically evaluating the physical safety of LLMs in embodied decision making.
Outcome: The proposed framework assesses the physical safety of LLMs in embodied decision making.
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Despite advances in artificial intelligence, building social intelligence remains a challenge.
Approach: They propose a task to explain why people laugh in a video and a dataset to do this.
Outcome: The proposed dataset generates plausible explanations for laughter in video and in-the-wild videos.
Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for conversational recommender systems lack specific user preferences and explanations for recommendations . current datasets lack specific preferences, hindering high-quality recommendations despite advances in large language models .
Approach: They propose to synthesize a conversational recommendation dataset with persona- and knowledge-augmented LLM simulators to address these challenges.
Outcome: The proposed dataset outperforms baselines in human and automatic evaluations.
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches.
Approach: They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought .
Outcome: The proposed system outperforms text-based models in singability and contextual accuracy.
Representation Bending for Large Language Model Safety (2025.acl-long)

Copied to clipboard

Challenge: Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and fail to generalize across unseen attacks.
Approach: They propose a new approach that disrupts representations underlying harmful behaviors in Large Language Models by using loss-based fine-tuning.
Outcome: The proposed approach outperforms existing methods such as Circuit Breaker, RMU, and NPO with 95% reduction in attack success rates across diverse jailbreak benchmarks.
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for generating speech from facial images rely on pre-trained visual encoders and fine-tune them to align with speech embeddings.
Approach: They propose to derive corresponding voices from facial images using face-to-voice synthesis, which derives corresponding voice from facial image.
Outcome: The proposed approach significantly improves face-voice congruence and synthesis stability.
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Visual arguments rely on images to persuade viewers to do or believe something .
Approach: They propose three tasks for evaluating visual argument understanding . they use visual premises, commonsense premises and reasoning trees to analyze visual arguments .
Outcome: The proposed tasks evaluate visual argument understanding using a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments.
Aligning Large Language Models by On-Policy Self-Judgment (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches for aligning large language models with human preferences face a trade-off that requires a separate reward model for on-policy learning.
Approach: They propose a new alignment framework that does on-policy learning and is parameter efficient . they propose Judge-augmented Supervised Fine-Tuning to train a single model to act as a policy and a judge.
Outcome: The proposed framework outperforms baselines in preference benchmarks and rejecting sampling by itself improves performance without additional evaluator.
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding (2026.acl-long)

Copied to clipboard

Challenge: Recent studies focus on surface-level features, overlooking how design choices influence user behavior at scale.
Approach: They propose a benchmark for multimodal understanding of how UI/UX design affects user behavior built on 300 real-world UI image pairs from industry A/B tests.
Outcome: The proposed benchmarks show that models exhibit limited understanding of the behavioral impact of UI/UX design.
Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal UNcommonsense (MUN) is a benchmark designed to evaluate models’ ability to handle scenarios that deviate from typical visual or contextual expectations.
Approach: They propose a retrieval-based in-context learning framework that transfers reasoning capabilities from larger models to smaller ones without additional training.
Outcome: The proposed method improves on baseline ICL methods by 8.3% over previous methods.
Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use large language models are not available due to ethical concerns, and data privacy concerns are a concern.
Approach: They propose a multi-turn dialogue dataset that emulates real-life counseling interactions using the goal-oriented approach of Cognitive Behavioral Therapy (CBT).
Outcome: The proposed model outperforms other models in counseling skills, highlighting its effectiveness and potential as a counseling agent.
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) absorb social and cultural biases embedded in vast web-scale corpora and are increasingly deployed in high-stakes domains such as hiring, education, and law.
Approach: They propose a framework to investigate counterfactual unfairness through humor by observing how the model’s responses change when we swap who speaks and who is addressed while holding other factors constant.
Outcome: The proposed framework covers humor generation refusal, speaker intention inference, and relational/societal impact prediction tasks.
SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization (2023.emnlp-main)

Copied to clipboard

Challenge: a dataset of 1.5 million conversations distilled from everyday spoken situations is limited in scale due to its associated costs.
Approach: They propose to make SODA a publicly available, million-scale high-quality social dialogue dataset . they contextualize social commonsense knowledge from a knowledge graph to distill broad spectrum of social interactions .
Outcome: The proposed dataset is the first publicly available, million-scale high-quality social dialogue dataset.
ProsocialDialog: A Prosocial Backbone for Conversational Agents (2022.emnlp-main)

Copied to clipboard

Challenge: Existing dialogue systems fail to respond properly to potentially unsafe user utterances . existing systems either ignore or passively agree with unsafe content .
Approach: They introduce a dataset to teach conversational agents to respond to problematic content following social norms.
Outcome: The proposed dataset shows that ProsocialDialog generates more socially acceptable dialogues than existing models.
Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms (2023.emnlp-main)

Copied to clipboard

Challenge: NormLens is a visual-grounded framework for understanding commonsense norms . state-of-the-art models are not well-aligned with human annotation, we show .
Approach: They propose a visual-grounded framework to study commonsense norms by NormLens . they find that models are not well-aligned with human annotation .
Outcome: The proposed model judgments and explanations are not well-aligned with human annotations.
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis.
Approach: They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Outcome: The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues (2025.acl-long)

Copied to clipboard

Challenge: Existing large language models fail to incorporate nonverbal elements into conversational experiences.
Approach: They propose a multimodal language model that generates nonverbal cues alongside text . their dataset is annotated with time-aligned text, facial expressions, and body language .
Outcome: The proposed model generates nonverbal languages and text, corresponding to conversational input.
Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents (2023.emnlp-main)

Copied to clipboard

Challenge: a human-like chatbot requires commonsense reasoning to comprehend and respond to information . however, identifying and aggregating key evidence within a single hop is a challenge . a knowledge distillation framework is proposed that leverages LLMs as unreliable teachers .
Approach: They propose a framework that leverages large language models as unreliable teachers to facilitate multi-hop reasoning over a dialogue context.
Outcome: The proposed framework leverages LLMs as unreliable teachers and selectively distills consistent and helpful rationales via alignment filters.
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models (MLLMs) offer new opportunities for higher-level scene understanding, but they require labor-intensive, expert annotation.
Approach: They propose a dataset that combines 2K human-verified images with 22K image-description pairs to provide a more accurate representation of pedestrian scenes.
Outcome: The proposed dataset improves scalability while maintaining quality.
Tracing Mathematical Proficiency Through Problem-Solving Processes (2026.findings-acl)

Copied to clipboard

Challenge: Knowledge Tracing (KT) models a learner's evolving knowledge state over time, but lacks the rich information embedded in students' problem-solving processes.
Approach: They propose a framework that uses a teacher-student-teacher pipeline to extract students’ Mathematical Proficiency (MP) as intermediate representation.
Outcome: The proposed framework improves the prediction performance of existing KT methods and provides interpretable explanations by explicitly modeling students’ mathematical proficiency.
C2: Scalable Auto-Feedback for LLM-based Chart Generation (2025.naacl-long)

Copied to clipboard

Challenge: generating high-quality charts with Large Language Models presents significant challenges due to limited data and the high cost of curation.
Approach: They propose a referencefree automatic feedback generator to generate high-quality charts with Large Language Models.
Outcome: The proposed framework outperforms baselines and shows that it significantly improves data diversity.
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Prior work has used LLMs to generate programming language and applied external compilers for such tasks.
Approach: They propose a framework that expresses task-level logic with pseudocode and tailors it to each instance and simulates execution of it.
Outcome: The proposed framework outperforms baselines in diverse reasoning tasks.
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have led to their adaptation as conversational agents.
Approach: They propose a new benchmark that uses 8K multi-choice questions to assess the personality of Large Language Models.
Outcome: The proposed personality test outperforms existing personality tests for LLMs in reliability and validity.
DUSK: Do Not Unlearn Shared Knowledge (2026.findings-acl)

Copied to clipboard

Challenge: Recent work suggests that machine learning models are indistinguishable from models trained on retain sets.
Approach: They propose a benchmark to evaluate machine unlearning under realistic knowledge overlap . they construct documents containing both shared and unique knowledge .
Outcome: The proposed model is indistinguishable from a model retrained on the retain set while only forget-specific content is removed.
Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-Speech (2025.findings-naacl)

Copied to clipboard

Challenge: Emotional dialogue speech synthesis (EDSS) aims to generate expressive speech by leveraging the dialogue context between interlocutors.
Approach: They propose a large language model to generate holistic emotion tags based on prior dialogue context and pinpoint key words in the target utterance that align with the predicted emotional state.
Outcome: The proposed method improves emotional expressiveness and facilitates automatic emotion speech generation during inference.
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models struggle when faced with unseen domains or languages.
Approach: They propose a framework that leverages the broad knowledge of an MLLM to generate cross-modal pre-questions (preQs) before retrieval.
Outcome: Experiments show that PREMIR outperforms existing retrievers on out-of-distribution benchmarks, including closed-domain and multilingual settings, outperforming strong baselines across all metrics.
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on embodied agents have addressed the importance of exploration in environments where tasks and solutions are not predefined.
Approach: They propose a virtual escape room that evaluates AI models in a dynamic environment . they propose to integrate memory management and reasoning into the simulation .
Outcome: The proposed model improves in dynamic and exploration-driven environments by integrating memory management and reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations