Papers by Hanseok Oh

6 papers
Generative Multi-hop Retrieval (2022.emnlp-main)

Copied to clipboard

Challenge: A bi-encoder approach to text retrieval has limitations in multi-hop settings; the reformulated query gets longer as the number of hops increases, which further tightens the embedding bottleneck of the query vector.
Approach: They propose an encoder-decoder model that performs multi-hop retrieval by simply generating the entire text sequences of the retrieval targets.
Outcome: The proposed model achieves comparable or higher performance than bi-encoder models in five datasets while demonstrating superior GPU memory and storage footprint.
Subject-level Inference for Realistic Text Anonymization Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing text anonymization evaluations assume only a single data subject, ignoring multi-subject scenarios.
Approach: They propose a benchmark that shifts the unit of evaluation from text spans to individuals . they show that subject-level inference protection drops as low as 33% when masked .
Outcome: The proposed benchmark reduces the amount of protection available when PII spans are masked.
Nonparametric Decoding for Generative Retrieval (2023.findings-acl)

Copied to clipboard

Challenge: Existing text retrieval models depend on the information encoded in its parameters without external memory, its information capacity is limited and fixed.
Approach: They propose a nonparametric decoding approach which uses external memory instead of vanilla vocab embeddings as decoder voka embedds.
Outcome: The proposed model can utilize parametric and nonparametric space.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
KTRL+F: Knowledge-Augmented In-Document Search (2024.naacl-long)

Copied to clipboard

Challenge: KTRL+F is a knowledge-augmented in-document search that requires real-time identification of all semantic targets within a document with the awareness of external sources through a single natural query.
Approach: They propose a knowledge-augmented in-document search that requires real-time identification of all semantic targets within a document with the awareness of external sources through a single natural query.
Outcome: The proposed model reduces time for searching with less queries and reduced extra visits to other sources for collecting evidence.
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact (2026.eacl-long)

Copied to clipboard

Challenge: Prior work on instruction tuning datasets combined these data types without examining their distinct effects.
Approach: They investigate how training LLMs with or without context affects model behavior and performance . they find that using context-augmented data as the backbone for vision-language models reduces hallucination .
Outcome: The proposed training with context-augmented data reduces hallucination and improves grounding in the visual domain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations