Papers by Hyounghun Kim
Continuous Language Generative Flow (2021.acl-long)
Copied to clipboard
| Challenge: | Recent years have witnessed various types of generative models for natural language generation (NLG), especially RNNs or transformers. |
| Approach: | They propose a flow-based language generation model that adapts flow-derived generative models to language generation via continuous input embeddings, adapted affine coupling structures, and a novel architecture for autoregressive text generation. |
| Outcome: | The proposed model improves on QG and NMT and improves performance over baselines on SQuAD and TVQA and NML16. |
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items? (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for estimating the cognitive complexity of reading comprehension items are expensive, time-consuming, and subject to rater variability. |
| Approach: | They propose to use two dimensions to estimate cognitive complexity of RC items to focus on evidence Scope and transformation level to estimate the cognitive complexity. |
| Outcome: | The proposed models can estimate the cognitive complexity of items by focusing on two dimensions—Evidence Scope and Transformation Level—that indicate the degree of cognitive burden involved in reasoning about the answer. |
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition (2026.acl-long)
Copied to clipboard
| Challenge: | Accented speech remains a persistent challenge for automatic speech recognition (ASR) Accent-agnostic approaches improve robustness but struggle with heavily accented or unseen varieties . |
| Approach: | They propose a Mixture-of-Experts architecture with intermediate CTC supervision that promotes expert specialization and generalization. |
| Outcome: | Experiments show that the proposed architecture improves on accented speech . the proposed framework is based on a mixture-of-experts architecture with intermediate supervision . |
Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations (2023.emnlp-main)
Copied to clipboard
| Challenge: | open-domain chatbots focus on short single-session dialogue, neglecting the potential need for understanding contextual information in multiple consecutive sessions. |
| Approach: | They propose a 1M multi-session dialogue dataset for integrating time intervals and speaker relationships into a long-term conversation setup. |
| Outcome: | The proposed model can generate coherent responses according to time intervals and speaker relationships with high user engagement without contradiction in a long-term conversation setup. |
Improving Visual Question Answering by Referring to Generated Paragraph Captions (P19-1)
Copied to clipboard
| Challenge: | Empirical results show that paragraph captions help answer more visual questions . |
| Approach: | They propose a visual and textual question answering model which uses paragraph captions as input . they use cross-attention to extract related information, then consensus to fuse the inputs . |
| Outcome: | Empirical results show that paragraph captions help answer more visual questions . the proposed model significantly improves the baseline model . |
Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA (2020.acl-main)
Copied to clipboard
| Challenge: | Recent years have witnessed a paradigm shift in the way we get our information, and a lot of it. |
| Approach: | They propose a video question answering model which integrates multi-modal input sources and finds temporally relevant information to answer questions. |
| Outcome: | The proposed model outperforms the state-of-the-art on a TVQA dataset. |
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models? (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent reasoning language models (RLMs) achieve strong performance on complex reasoning tasks, yet they still exhibit a multilingual reasoning gap. |
| Approach: | They propose a strategy that incorporates an English translation into the initial reasoning trace when an understanding failure is detected. |
| Outcome: | The proposed strategy incorporates an English translation into the initial reasoning trace when an understanding failure is detected. |
PanicToCalm: A Proactive Counseling Agent for Panic Attacks (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for training such models are limited due to ethical and logistical issues. |
| Approach: | They propose a dataset that includes high-distress episodes constructed from first-person narratives and structured around the principles of Psychological First Aid. |
| Outcome: | The proposed model outperforms baseline models in counselor-side metrics and client affect improvement. |
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for generating speech from facial images rely on pre-trained visual encoders and fine-tune them to align with speech embeddings. |
| Approach: | They propose to derive corresponding voices from facial images using face-to-voice synthesis, which derives corresponding voice from facial image. |
| Outcome: | The proposed approach significantly improves face-voice congruence and synthesis stability. |
Self-Correcting Code Generation Using Small Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has demonstrated that self-correction is a powerful tool for code generation, but whether it is effective for smaller models remains unexplored. |
| Approach: | They propose a method that trains small language models to maintain correct outputs while progressively correcting incorrect outputs as turns proceed. |
| Outcome: | The proposed approach improves the ability of small language models for multi-turn code correction. |
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable performances in general domains and are now extending into the expert domain of law. |
| Approach: | They propose a Korean Benchmark for Legal EXplainable QA (KoBLEX) that evaluates provision-grounded, multi-hop legal reasoning. |
| Outcome: | The proposed method outperforms baselines and shows a high correlation with human judgments. |
CoSIm: Commonsense Reasoning for Counterfactual Scene Imagination (2022.naacl-main)
Copied to clipboard
| Challenge: | a new dataset evaluates the ability of AI systems to reason about scene change imagination . a large human-model performance gap exists in the dataset . |
| Approach: | They propose a dataset to evaluate AI's ability to reason about scene change imagination . they use an image and a commonsense question to imagine a counterfactual scene change . |
| Outcome: | The proposed dataset evaluates the ability of AI systems to reason about scene change imagination. |
RESIN-11: Schema-guided Event Prediction for 11 Newsworthy Scenarios (2022.naacl-demo)
Copied to clipboard
Xinya Du, Zixuan Zhang, Sha Li, Pengfei Yu, Hongwei Wang, Tuan Lai, Xudong Lin, Ziqi Wang, Iris Liu, Ben Zhou, Haoyang Wen, Manling Li, Darryl Hannan, Jie Lei, Hyounghun Kim, Rotem Dror, Haoyu Wang, Michael Regan, Qi Zeng, Qing Lyu, Charles Yu, Carl Edwards, Xiaomeng Jin, Yizhu Jiao, Ghazaleh Kazeminejad, Zhenhailong Wang, Chris Callison-Burch, Mohit Bansal, Carl Vondrick, Jiawei Han, Dan Roth, Shih-Fu Chang, Martha Palmer, Heng Ji
| Challenge: | Existing methods for event prediction are incomplete and noisy. |
| Approach: | They propose to use news-related event schemas to extract newsworthy events . they build a demo website and include a video demonstrating the framework . |
| Outcome: | The proposed framework can be applied to a wide variety of newsworthy scenarios. |
ChronoBias: A Benchmark for Evaluating Temporal Group Bias in the Time-sensitive Knowledge of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using a template-based semi-automated generation method, we evaluate time-conditional group bias in time-sensitive knowledge of large language models (LLMs). |
| Approach: | They propose a template-based semi-automated generation method to construct a time-conditional group bias benchmark. |
| Outcome: | The proposed method balancing quality-quantity trade-off in existing benchmark curation approaches. |
Sound of Story: Multi-modal Storytelling with Audio (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on storytelling with sound have focused on visuals and sounds, but little attention has been given to sound. |
| Approach: | They propose to establish a new component called background sound which is story context-based audio without any linguistic information. |
| Outcome: | The proposed dataset is the largest well-curated dataset for storytelling with sound . it contains 27,354 stories with 19.6 images per story and 984 hours of speech-decoupled audio . |
NDH-Full: Learning and Evaluating Navigational Agents on Full-Length Dialogue (2021.emnlp-main)
Copied to clipboard
| Challenge: | Vision-and-Dialogue Navigation is one of the tasks that evaluate the agent’s ability to interact with humans for assistance and navigate based on natural language responses. |
| Approach: | They propose a vision-and-dialogue navigation task which evaluates the agent's ability to interact with humans and navigate based on natural language responses. |
| Outcome: | The proposed model performs well on the Navigation from Dialogue History task, but it is not evaluated by the primary metric Goal Progress. |
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions (2025.acl-long)
Copied to clipboard
| Challenge: | Multimodality has been explored in multi-party and multi-session conversations, but task-specific constraints have hindered its seamless integration into dynamic, natural conversations. |
| Approach: | They propose a multimodal conversation dataset and a model with multimodal memory retrieval to equip chatbots with "eyes and ears" they aim to integrate multimodality into chatbot interactions by integrating visual and auditory inputs into the chatbot. |
| Outcome: | The proposed model demonstrates the ability to engage in long-term conversations with multiple speakers in complex, real-world-like settings, effectively processing visual and auditory inputs to understand and respond appropriately. |
Mixed-Session Conversation with Egocentric Memory (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent dialogue systems exhibit an inability to replicate dynamic, continuous, long-term interactions involving multiple partners. |
| Approach: | They propose a multi-session dialogue system that builds on real-world interactions by integrating deep layered interactions and widening conversation networks. |
| Outcome: | The proposed system is based on a dataset of 6 consecutive dialogue episodes with four speakers (one main speaker and three partners) appearing in each episode. |
Collective Critics for Creative Story Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have enabled the automatic generation of long-form stories containing several thousand words. |
| Approach: | They propose a framework that creates a story plan and generates based on it, and integrates 'creative' revision mechanism into long-form story generation process. |
| Outcome: | The proposed framework can significantly enhance story creativity and reader engagement while maintaining narrative coherence. |
Revealing the Inherent Instructability of Pre-Trained Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Pre-trained large language models perform multitask learning during their pre-training . a new technique, Response Tuning, removes the instruction and its corresponding mapping to the response from instruction tuning. |
| Approach: | They propose a method which removes the instruction and its mapping to the response from instruction tuning. |
| Outcome: | The proposed model can respond to a wide range of instructions . it can recognize and reject unsafe queries after learning from response data. |
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for difficulty-controlled reading comprehension item generation rely on a single agent prompting approach. |
| Approach: | They propose a multi-agent framework for Feature-constrained Item Generation where multiple LLM agents collaborate to generate and iteratively revise items based on intended constraints. |
| Outcome: | The proposed method generates items with monotonically increasing difficulty at higher rates than baselines. |
ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments (2020.findings-emnlp)
Copied to clipboard
| Challenge: | embodied agents are expected to perform specific tasks after reaching the destination . a novel vision-and-language navigation task is designed to support this task . |
| Approach: | They combine vision-and-language navigation, assembling objects and object referring expression comprehension to create a joint navigation-and assembly task. |
| Outcome: | The proposed task is based on vision-and-language navigation and assembly . it uses human-written navigation and assembling instructions and ground truth trajectories . the large model-human performance gap shows that the task is challenging and wide scope for future work. |
Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing (2026.findings-acl)
Copied to clipboard
| Challenge: | Knowledge Tracing (KT) aims to predict learners’ future performance from past interactions, but they overlook the procedural dynamics of problem solving. |
| Approach: | They propose a framework that enriches item representations by integrating dynamic procedural solution information. |
| Outcome: | Experiments on XES3G5M and NIPS34 show that BAIM outperforms strong pretraining-based baselines, achieving particularly large gains under repeated learner interactions. |