Papers by Junghwan Kim

7 papers
ABLE: Agency-BeLiefs Embedding to Address Stereotypical Bias through Awareness Instead of Obliviousness (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies in Natural Language Processing (NLP) have unveiled a concerning issue: stereotypical biases associated with demographic groups are prevalent.
Approach: They propose an approach that actively encodes stereotypical biases into the embedding space by integrating stereotypes into a model that acquires agency and belief scores rather than directly representing stereotypes.
Outcome: The proposed model can learn agency and belief stereotypes while preserving the language model’s proficiency.
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are susceptible to jailbreak prompts that can elicit harmful or inappropriate responses.
Approach: They propose a black-box framework for automated red teaming that integrates a Multi-Agent System with a Strategy-Response Multiplex Network and employs network-driven optimization to sample effective attack strategies.
Outcome: The proposed framework surpasses existing methods and achieves higher attack success rate (ASR) at lower computational cost.
Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains (2025.emnlp-main)

Copied to clipboard

Challenge: Authorship representation (AR) models capture an author's distinctive writing style by encoding documents written by the same author as nearby vectors in the embedding space.
Approach: They propose a method that incorporates probabilistic content masking and language-aware batching to improve contrastive learning by reducing cross-lingual interference.
Outcome: The proposed model outperforms monolingual baselines in 21 out of 22 non-English languages and reaches a maximum gain of 15.91% in a single language.
Interpreting Style Representations via Style-Eliciting Prompts (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has attempted to explain learning of style representations by generating natural language descriptions with large language models (LLMs) conditioned on input text.
Approach: They propose a framework for interpreting style representations through style-eliciting prompts by prompting an LLM to generate text conditioned on these features.
Outcome: The proposed framework outperforms baselines that directly prompt LLMs with target text, and achieves superior performance in both style description and style imitation.
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
Race, Gender, and Age Biases in Biomedical Masked Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models can be used to identify and eliminate healthcare disparities.
Approach: They examine social biases present in biomedical masked language models . they curate prompts based on evidence-based practice and compare generated diagnoses .
Outcome: The proposed models are less biased than BERT in gender, while the opposite is true for race and age.
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) must possess an understanding of the nation’s culture and basic knowledge.
Approach: They propose to construct a national alignment benchmark, KorNAT, which measures the alignment between an LLM and a targeted country from two perspectives: social value alignment and common knowledge alignment.
Outcome: The proposed model passes the national alignment score of 7 LLMs, indicating there is room for improvement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations