Papers by Seunghee Kim

5 papers
OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for multimodal large language models suffer from limitations . modality shortcuts and biased reasoning paths are common in such models .
Approach: a new benchmark evaluates omni-modal multi-hop reasoning using 6,144 questions . authors propose OMHBench to address these limitations by comparing modalities .
Outcome: OMHBench evaluates omni-modal multi-hop reasoning on 6,144 questions with balanced reasoning paths . evaluation of 13 state-of-the-art models shows performance gap exists between MLLMs and open-source models .
FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for multimodal large language models lack data contamination and complex queries . financial cross-modal multi-hop reasoning is difficult to evaluate and requires precise cross-module reasoning .
Approach: They propose a benchmark to analyze the reasoning capabilities of multimodal large language models.
Outcome: The proposed model is categorized into three difficulty levels—easy, medium, and hard—for step-by-step evaluation.
Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing (2024.findings-emnlp)

Copied to clipboard

Challenge: Visual speech processing requires context modeling due to the ambiguous nature of lip movements.
Approach: They propose a framework to maximize the context modeling capability by bringing the power of LLMs.
Outcome: The proposed framework maximizes the power of visual speech processing by bringing it to the forefront of the field.
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens (2026.acl-long)

Copied to clipboard

Challenge: Recent studies have explored token-wise loss regularizers that prioritize informative tokens, but rely on ground-truth confidence or external linguistic parsers, which limits their ability to capture contextual information or the model’s overall predictive state.
Approach: They propose an Entropy-guided Token Weighting (ETW) token-level unlearning regularizer that uses entropy of the predictive distribution as a proxy for token informativeness.
Outcome: The proposed token-level unlearning regularizer can achieve more effective unlearning while better preserving model utility than existing token-based approaches.
Constructing Korean Learners’ L2 Speech Corpus of Seven Languages for Automatic Pronunciation Assessment (2024.lrec-main)

Copied to clipboard

Challenge: Multilingual L2 speech corpora for automatic speech assessment are currently available, but lack comprehensive annotations of L2 from non-native speakers of various languages.
Approach: They propose to use Korean learners’ L2 speech corpus of seven languages to develop automatic speech assessment.
Outcome: The proposed corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each in French, German, Spanish, and Russian).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations