Papers by Yujin Kim

11 papers
Visual–Linguistic Abductive Reasoning with LLMs for Knowledge-based Visual Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Recent efforts to leverage large language models for reasoning focus on visual perception and language reasoning as separate processes.
Approach: They propose a method that integrates visual and linguistic modalities into interpretable abductive reasoning chains.
Outcome: The proposed method improves performance on AOKVQA, OKVQA and GQA by 2.31% . it uses fuzzy scoring to select the most coherent combination, enabling unified reasoning .
Self-Training Elicits Concise Reasoning in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Chain-of-thought reasoning has enabled large language models to use additional computation through intermediate tokens to solve complex tasks, but current models often generate more tokens than necessary to accomplish the task, incurring extraneous inference costs.
Approach: They propose to fine-tune models with self-generated concise reasoning paths obtained by best-of-N sampling and few-shot conditioning in task-specific settings to elicit concise reasoning.
Outcome: The proposed method reduces output tokens by 30% on GSM8K and MATH while maintaining average accuracy.
FinHarmBench: Financial Jailbreak Benchmark and Unsupervised Safety Fine-Tuning via Refusal Steering Distillation (2026.acl-industry)

Copied to clipboard

Challenge: Existing safety benchmarks focus on general harms and lack the granularity needed to capture domain-specific financial threats.
Approach: They propose a benchmark to evaluate financially harmful and confusable benign prompts.
Outcome: The proposed framework improves refusal behavior without annotating refusal responses.
Diagnosing Spatial Consistency across Perspectives and Viewpoints in Large Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing models assess spatial capabilities from a static, single-view and egocentric perspective, failing to capture the dynamic nature of real-world spatial cognition.
Approach: They propose a benchmark to diagnose spatial reasoning capabilities using a 360 field of view.
Outcome: The proposed benchmark evaluates allocentric and egocentric reasoning capabilities from multiple perspectives in high-quality 3D environments.
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to creating inclusive vision-language models rely on human annotators, making it labor-intensive and creating cognitive burdens.
Approach: They propose a semi-automated framework for constructing cultural VLM benchmarks . they use an annotated sample of Korean culture to generate questions .
Outcome: The proposed framework is based on a Korean culture dataset and shows that open-source models lag behind proprietary ones in understanding Korean culture.
Hidden Persuaders: LLMs’ Political Leaning and Their Influence on Voters (2024.emnlp-main)

Copied to clipboard

Challenge: This paper examines the political leanings of large language models (LLMs) in the 2024 election.
Approach: They propose to use large language models to examine users' political leanings in the 2024 presidential election to determine their political preference.
Outcome: The proposed models show that they have a political leaning and can influence political views in the 2024 presidential election.
HARE: Explainable Hate Speech Detection with Step-by-Step Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent benchmarks have attempted to identify and explain hate speech but lack the reasoning to supervise detection models.
Approach: They propose a framework that uses large language models to fill in the gaps in hate speech explanations by using existing annotations.
Outcome: The proposed framework outperforms baselines on SBIC and Implicit Hate using model-generated data and improves generalization to unseen datasets.
NASH: A Simple Unified Framework of Structured Pruning for Accelerating Encoder-Decoder Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Structured pruning methods have proven effective in reducing the model size and accelerating inference speed in various network architectures.
Approach: They propose a framework that narrows the encoder and shortens the decoder networks of encoder-decoder models.
Outcome: The proposed framework reduces the number of decoder layers and improves generation quality.
Carpe diem: On the Evaluation of World Knowledge in Lifelong Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Current language models are trained on static data, implying that the encoded knowledge could go wrong as time passes.
Approach: They propose a temporally evolving question-answering benchmark for language models . they use Wikipedia databases to test language models for dynamic knowledge in ever-changing world .
Outcome: The proposed task aims to model the evolution-adaptability of language models in the real world.
Towards Formality-Aware Neural Machine Translation by Leveraging Context Information (2023.findings-emnlp)

Copied to clipboard

Challenge: Formality is one of the most important linguistic properties to determine the naturalness of translation.
Approach: They propose a method to explicitly inform neural machine translation models by pinpointing key informative tokens using a formality classifier.
Outcome: The proposed method improves translation quality and conforms to the appropriate syntax.
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to align Large Language Models with human preferences fail to maintain general knowledge and alignment when faced with personalized preferences.
Approach: They propose a method that utilizes the initial responses of the reference model to mitigate forgetting while accommodating personalized alignment.
Outcome: The proposed approach mitigates forgetting while accommodating personalized alignment while preserving global knowledge and general alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations