Papers by Hyukhun Koh

10 papers
Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Gender bias is a significant issue in machine translation, but most studies focus on debiasing bilingual models without consideration for multilingual systems.
Approach: They propose a method which debiases bilingual models for unambiguous cases where there is a single correct translation.
Outcome: The proposed method improves gender accuracy by a wide margin without hampering translation performance.
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities but their misuse for harmful purposes remains a concern.
Approach: They propose a jailbreaking technique that exploits weaknesses in LLMs' architecture . they propose abductive framing and symbolic encoding to bypass safeguards .
Outcome: The proposed technique achieves over 95% attack success rate on GPT-series models and 70% across all targets.
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing toxicity metrics rely on encoder models trained on specific toxicity datasets, which are susceptible to out-of-distribution (OOD) problems and depend on the dataset’s definition of toxicity.
Approach: They propose a robust metric grounded on LLMs to flexibly measure toxicity according to the given definition by analysing toxicity factors and intrinsic toxic attributes.
Outcome: The proposed metric improves on conventional metrics by 12 points in the F1 score and shows that upstream toxicity significantly influences downstream metrics, suggesting that LLMs are unsuitable for toxicity evaluations within unverified factors.
VLind-Bench: Measuring Language Priors in Large Vision-Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Large Vision-Language Models suffer from a problem known as language prior . such language priors can lead to undesirable biases and hallucinations when dealing with images that are out of distribution.
Approach: They propose a benchmark to measure the language priors of Large Vision-Language Models.
Outcome: The proposed benchmark is the first specifically designed to measure the language priors, or blindness, of LVLMs.
Generating Diverse Hypotheses for Inductive Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies suggest that large language models (LLMs) can engage in inductive reasoning by sampling multiple hypotheses about the rules and selecting the one that best explains the observations.
Approach: They propose to increase the temperature parameter to enhance diversity by sampling multiple hypotheses and selecting the one that best explains the observations.
Outcome: The proposed method improves diversity while maintaining text quality while increasing temperature.
Public Data Assisted Differentially Private In-Context Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: In-context learning has shown remarkable performance across tasks without fine-tuning . however, recent studies have highlighted the risk of private data leakage through the prompt in ICL .
Approach: They propose a private in-context learning algorithm that effectively balances privacy protection and model utility.
Outcome: The proposed algorithm is robust against membership inference attacks and is robust to membership infertility attacks.
Evaluating Visual Narrative Coherence in Story Visualization via Diversified Storylines (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation metrics and datasets often neglect visual continuity and narrative diversity.
Approach: They propose a visual context-aware metric for story visualization that uses large vision-language models to jointly assess caption fidelity and inter-image consistency.
Outcome: The proposed framework achieves a Spearman’s correlation comparable to human agreement on two benchmarks and blends diverse and controlled narrative elements at adjustable ratios, producing challenging evaluation sets.
DPP-TTS: Diversifying prosodic features of speech via determinantal point processes (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep generative models have succeeded in synthesizing human-like speech.
Approach: They propose a text-to-speech model with a prosody diversifying module that considers perceptual diversity in each sample and among multiple samples.
Outcome: The proposed model generates speech samples with more diversified prosody than baselines in the side-by-side comparison test considering the naturalness of speech at the same time.
Fine-grained Gender Control in Machine Translation with Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing work on controlled translation has only considered a simplified setup of one target gender for input.
Approach: They propose a Gender-of-Entity prompting method for machine translation that takes the gender of the ambiguous entity as additional input and propose to use it to translate with correct gender inflections.
Outcome: The proposed method instructs the model with fine-grained entity-level gender information to translate with correct gender inflections.
Conditional [MASK] Discrete Diffusion Language Model (2025.emnlp-main)

Copied to clipboard

Challenge: Auto-regressive models excel in natural language processing but struggle to generate diverse text and lack controllability.
Approach: They propose entropy-adaptive Gibbs sampling and entropic-based noise scheduling to counterbalance each model’s shortcomings.
Outcome: The proposed framework outperforms baseline models and achieves the best quality-diversity tradeoff, demonstrating its effectiveness in non-autoregressive text generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations