Papers by Junhyeok Kim

4 papers
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild (2025.findings-naacl)

Copied to clipboard

Challenge: EgoSpeak predicts when an agent should begin speaking based on egocentric streaming video.
Approach: They propose a framework for real-time speech initiation prediction in egocentric streaming video by modeling the conversation from the camera wearer's first-person perspective.
Outcome: The proposed framework outperforms random and silence-based baselines in real time and highlights the importance of multimodal input and context length in effectively deciding when to speak.
Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms (2023.emnlp-main)

Copied to clipboard

Challenge: NormLens is a visual-grounded framework for understanding commonsense norms . state-of-the-art models are not well-aligned with human annotation, we show .
Approach: They propose a visual-grounded framework to study commonsense norms by NormLens . they find that models are not well-aligned with human annotation .
Outcome: The proposed model judgments and explanations are not well-aligned with human annotations.
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues (2025.acl-long)

Copied to clipboard

Challenge: Existing large language models fail to incorporate nonverbal elements into conversational experiences.
Approach: They propose a multimodal language model that generates nonverbal cues alongside text . their dataset is annotated with time-aligned text, facial expressions, and body language .
Outcome: The proposed model generates nonverbal languages and text, corresponding to conversational input.
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models (MLLMs) offer new opportunities for higher-level scene understanding, but they require labor-intensive, expert annotation.
Approach: They propose a dataset that combines 2K human-verified images with 22K image-description pairs to provide a more accurate representation of pedestrian scenes.
Outcome: The proposed dataset improves scalability while maintaining quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations