Papers by Sean Kim

9 papers
Evaluating Language Models as Synthetic Data Generators (2025.acl-long)

Copied to clipboard

Challenge: Prior studies have focused on developing effective data generation methods, but lack systematic comparison of different LMs as data generators in a unified setting.
Approach: They propose to use a benchmark to compare language models' data generation abilities against a set of standardized settings and metrics.
Outcome: The proposed benchmark provides standardized settings and metrics to evaluate LMs’ data generation abilities.
Scaling Evaluation-Time Compute with Reasoning Models as Evaluators (2026.findings-acl)

Copied to clipboard

Challenge: Language model (LM) evaluators that generate chain-of-thought reasoning are widely used for the assessment of LM responses.
Approach: They investigate whether increasing LMs' "thinking" time through scaling test-time compute can improve an LM's evaluation capability.
Outcome: The proposed reasoning models improve evaluation performance monotonically with the number of reasoning tokens generated, mirroring trends seen in LM reasoning.
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing open-source evaluation paradigms lack flexibility and performance . language model-based evaluation is cheap and scalable, but it is difficult to evaluate .
Approach: They propose a language model-based evaluation paradigm that uses a scalar indicator of quality to assess LM outputs.
Outcome: The proposed language model-based evaluation model is more powerful than its predecessor.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
Consistency of a Recurrent Language Model With Respect to Incomplete Decoding (2020.emnlp-main)

Copied to clipboard

Challenge: Neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition.
Approach: They propose to use a recurrent language model to address inconsistency in decoding algorithms that are inconsistent despite the fact that recursive language models are trained to produce sequences of finite length.
Outcome: The proposed methods prevent inconsistency in the proposed models.
Can Language Models be Biomedical Knowledge Bases? (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on probing LMs in the general domain but little attention has been given to whether they can be used as domain knowledge bases.
Approach: They propose to use 49K biomedical factual knowledge triples to probe LMs for biomedically . they find that biomedic LM can achieve up to 18.51% Acc@5 on retrieving biomedcial knowledge.
Outcome: The proposed biomedical factual knowledge probing benchmark achieves 18.51% Acc@5 on biomedically-relevant knowledge retrieval.
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored.
Approach: They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities.
Outcome: The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions.
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Large language models exhibit cultural and geopolitical biases when their outputs shape public opinion or reinforce dominant narratives.
Approach: They define two types of bias in large language models: model bias and inference bias through a two-phase evaluation.
Outcome: The proposed framework evaluates large language models on factual and disputable questions across four languages and question types.
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities (2026.acl-long)

Copied to clipboard

Challenge: Uncertainty quantification (UQ) for large language models is a key building block for daily applications.
Approach: They propose a general formulation of agent UQ that subsumes broad classes of existing UQ setups.
Outcome: The proposed framework is based on the first general formulation of agent UQ that subsumes broad classes of existing setups.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations