Papers by Jaewoo Kang

33 papers
FaVIQ: FAct Verification from Information-seeking Questions (2022.acl-long)

Copied to clipboard

Challenge: Existing fact verification datasets with crowdsourced claims introduce subtle biases that are difficult to control for.
Approach: They construct a large-scale fact verification dataset with ambiguous questions . they use a corpus of 188k claims to construct false and true claims .
Outcome: The proposed dataset outperforms models trained on the dataset FEVER or in-domain data by up to 17% absolute.
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Quantization is a practical solution for deploying Large Language Models in resource-constrained environments.
Approach: They propose an outlier-safe pre-training approach that prevents outlier formation . they validate a 1.4B-parameter model on 1 trillion tokens with no outliers .
Outcome: The proposed model achieves a 35.7 average score on 1 trillion tokens with 2% training overhead.
Generating Information-Seeking Conversations from Unlabeled Documents (2022.emnlp-main)

Copied to clipboard

Challenge: a novel framework for conversational question answering from unlabeled documents has been proposed . a large-scale dataset of synthetic conversations is available for use in real-world applications .
Approach: They propose a framework for conversational question answering from unlabeled documents . they propose 'SimSeek' framework that simulates conversation from unlabelled documents based on two scenarios .
Outcome: The proposed framework achieves state-of-the-art performance on a recent CQA benchmark, QuAC.
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) struggle with hallucinations and outdated knowledge.
Approach: They propose a retrieval-augmented generation framework for enhancing the reliability of RAG in biomedical contexts.
Outcome: The proposed framework outperforms the previous best medical RAG model by up to 5.6% across three medical question-answering benchmarks.
Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations (2023.acl-long)

Copied to clipboard

Challenge: Named entity recognition models rely on domain-specific dictionaries provided by experts . however, such dictionary sets are infeasible in many domains where they do not exist .
Approach: They propose a framework that generates NER datasets with high-coverage pseudo-dictionaries . phrase retrieval models are used to retrieve popular entities rather than rare ones .
Outcome: The proposed framework outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.
Benchmarking Direct Preference Optimization for Medical Large Vision–Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large vision-language models (LVLMs) are gaining traction in clinical tasks such as diagnostic support, report generation, and medical question answering.
Approach: They present a systematic evaluation of nine DPO variants applied to two leading medical LVLMs.
Outcome: The proposed model improves alignment and reduces severe hallucinations, but yields inconsistent gains over supervised fine-tuning.
ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluation methods do not assess whether large language models fully utilize contextual information.
Approach: They introduce a new metric to assess LLMs' ability to fully utilize contextual information.
Outcome: The proposed benchmark comprises 1,986 test instances spanning four long-context tasks with high IC scores in the domains of books, debates, medicine, and law.
Ranking Paragraphs for Improving Answer Recall in Open-Domain Question Answering (D18-1)

Copied to clipboard

Challenge: Recent work has combined open-domain question answering with machine comprehension models to find answers in a large knowledge source.
Approach: They propose a machine comprehension model that ranks paragraphs of retrieved documents for a higher answer recall with less noise.
Outcome: The proposed model improves on four open-domain QA datasets by 7.8% on average.
Learning Dense Representations of Phrases at Scale (2021.acl-long)

Copied to clipboard

Challenge: Existing phrase retrieval models rely on sparse representations and still underperform retriever-reader approaches.
Approach: They propose a method to learn phrase representations from reading comprehension tasks using negative sampling methods.
Outcome: The proposed model improves over previous models by 15%-25% absolute accuracy and matches the performance of state-of-the-art retrieval models.
Learn to Resolve Conversational Dependency: A Consistency Training Framework for Conversational Question Answering (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches do not explicitly train QA models on how to resolve conversational dependency, and thus these models are limited in understanding human dialogues.
Approach: They propose a framework that generates self-contained questions that can be understood without the conversation history and then trains a QA model with the pairs of original and self-constructed questions using a consistency-based regularizer.
Outcome: The proposed framework improves the models’ performance by up to 1.2 F1 on QuAC, and 5.2 F1 for CANARD, while addressing the limitations of the existing approaches.
CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents (2026.findings-acl)

Copied to clipboard

Challenge: Large language model agents rely on external memory to support knowledge reuse and reasoning tasks.
Approach: They propose a CLustering-based AGentic memory framework where an agent actively organizes memory . they employ an SLM-agent driven router to assign each new memory to a semantically coherent cluster .
Outcome: The proposed framework improves answer quality and robustness over previous memory systems.
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have enabled molecular reasoning for property prediction. however, toxicity arises from complex biological mechanisms, necessitating mechanistic reasoning for reliable prediction.
Approach: They propose a benchmark that evaluates organ-level toxicity reasoning across multiple organs . they find strong predictive performance does not necessarily imply reliable reasoning .
Outcome: The proposed benchmark evaluates toxicity prediction performance and reasoning quality across LLMs.
Biomedical NER for the Enterprise with Distillated BERN2 and the Kazu Framework (2022.emnlp-industry)

Copied to clipboard

Challenge: a low overall error rate is unavoidable in the field of bioNLP, argues a new study.
Approach: They propose a highly extensible open source framework that supports BioNLP for the pharmaceutical sector.
Outcome: The proposed framework is extensible and scalable and can support BioNLP for the pharmaceutical sector.
CompAct: Compressing Retrieved Documents Actively for Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to condense extensive documents with no loss of information are difficult to implement in real-world scenarios.
Approach: They propose a framework that employs an active strategy to condense extensive documents without losing key information.
Outcome: The proposed framework improves performance and compression rate on multi-hop question-answering benchmarks.
Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: In open-domain question answering, users often ask ambiguous questions (AQs) . one approach is to identify all possible interpretations of the AQ and generate a long-form answer addressing them all.
Approach: They propose a framework that generates a long-form answer addressing all possible interpretations of an ambiguous question.
Outcome: The proposed framework outperforms baselines on ASQA in a few-shot setup across metrics while surpassing fully-supervised baselines trained on the whole training set in terms of Disambig-F1 and Disambigo-ROUGE.
Consistency Training with Virtual Adversarial Discrete Perturbation (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for regularizing a model are agnostic to the training model and may not be effective for perturbed inputs.
Approach: They propose an augmentation method of adding a discrete noise that would incur the highest divergence between predictions by replacing tokens while keeping original semantics.
Outcome: The proposed method outperforms baselines on semi-supervised text classification tasks and a robustness benchmark.
Learning from Negative Samples in Biomedical Generative Entity Linking (2025.findings-acl)

Copied to clipboard

Challenge: Generative models are usually trained only with positive samples and do not explicitly learn from hard negative samples, which are entities that look similar but have different meanings.
Approach: They propose a framework that trains generative BioEL models using negative samples to learn from hard negative samples.
Outcome: The proposed framework outperforms baseline models by up to an average top-1 accuracy of 1.4% on five benchmarks.
CookingSense: A Culinary Knowledgebase with Multidisciplinary Assertions (2024.lrec-main)

Copied to clipboard

Challenge: CookingSense is a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes.
Approach: They introduce CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes.
Outcome: The proposed system improves retrieval augmented language models and food decision support systems.
Adversarial Subword Regularization for Robust Neural Machine Translation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for segmenting words into subword units are not robust enough to handle multiple subword candidates.
Approach: They propose to regularize subword segmentations that maximize the translation loss by using gradient signals during training to prevent erroneous segmentations of unseen words.
Outcome: The proposed method improves the performance of NMT models on low-resource and out-domain datasets.
Simple Questions Generate Named Entity Recognition Datasets (2022.emnlp-main)

Copied to clipboard

Challenge: Recent named entity recognition models rely on human-annotated datasets . however, in-domain dictionaries and sentences are often unavailable or expensive to construct for many entity types.
Approach: They propose an ask-to-generate approach which automatically generates NER datasets by asking natural language questions to an open-domain question answering system.
Outcome: The proposed model outperforms the previous best model by 19.5 F1 score on six benchmarks and achieves state-of-the-art performance.
Assessing LLM Reasoning Steps via Principal Knowledge Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Step-by-step reasoning has become a standard approach for large language models to tackle complex tasks.
Approach: They propose a framework that assesses the knowledge grounding of intermediate reasoning by using a large-scale repository of atomic knowledge essential for reasoning.
Outcome: The evaluation suite identifies missing or misapplied knowledge elements and provides crucial insights for uncovering fundamental reasoning deficiencies in LLMs.
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing (2024.findings-eacl)

Copied to clipboard

Challenge: Contrastive language-image pre-training models have demonstrated considerable success across various vision-language tasks, such as text-to-image retrieval.
Approach: They propose a fine-tuning approach to enhance the representations of CLIP models for paraphrases by leveraging large language models.
Outcome: The proposed model improves on baseline models across paraphrased retrieval, visual genome relation and attribution, and seven semantic textual similarity tasks.
Tracing Mathematical Proficiency Through Problem-Solving Processes (2026.findings-acl)

Copied to clipboard

Challenge: Knowledge Tracing (KT) models a learner's evolving knowledge state over time, but lacks the rich information embedded in students' problem-solving processes.
Approach: They propose a framework that uses a teacher-student-teacher pipeline to extract students’ Mathematical Proficiency (MP) as intermediate representation.
Outcome: The proposed framework improves the prediction performance of existing KT methods and provides interpretable explanations by explicitly modeling students’ mathematical proficiency.
“Killing Me” Is Not a Spoiler: Spoiler Detection Model using Graph Neural Networks with Dependency Relation-Aware Attention Mechanism (2021.eacl-main)

Copied to clipboard

Challenge: Several attention-based spoiler detection models are insufficient for utilizing dependency relations between context words.
Approach: They propose a new spoiler detection model called SDGNN that uses syntax-aware graph neural networks to detect dependency relations between context words.
Outcome: The proposed model outperforms existing models on two real-world benchmark datasets.
SCRIPTMIND: Crime Script Inference and Cognitive Evaluation for LLM-based Social Engineering Scam Detection System (2026.eacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promise in identifying deception, but their cognitive assistance potential remains underexplored.
Approach: They propose a framework for LLM-based scam detection that bridges automated reasoning and human cognition.
Outcome: The proposed framework outperforms GPT-4o in the Korean scam detection and phone scam simulations.
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct reasoning errors at specific steps of the reasoning process.
Approach: They propose a process reward modeling framework that leverages retrieval-augmented generation to verify each reasoning step against established medical knowledge bases.
Outcome: The proposed model improves on five medical QA benchmarks and two open-ended diagnostic tasks by 13.50% on MedQA.
Contextualized Sparse Representations for Real-Time Open-Domain Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing phrase retrieval models suffer from low accuracy due to limited scalability and speed . 'Open-domain question answering' is a task of answering generic factoid questions by looking up a large knowledge source, typically unstructured text corpora such as Wikipedia.
Approach: They aim to augment existing phrase retrieval models with contextualized sparse representations to improve the quality of each phrase embedding.
Outcome: The proposed model improves CuratedTREC and SQuAD-Open by 4% and 45x faster inference speeds over the existing model.
Can Language Models be Biomedical Knowledge Bases? (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on probing LMs in the general domain but little attention has been given to whether they can be used as domain knowledge bases.
Approach: They propose to use 49K biomedical factual knowledge triples to probe LMs for biomedically . they find that biomedic LM can achieve up to 18.51% Acc@5 on retrieving biomedcial knowledge.
Outcome: The proposed biomedical factual knowledge probing benchmark achieves 18.51% Acc@5 on biomedically-relevant knowledge retrieval.
Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to perform conversational search are sub-optimal due to the limited ability to incorporate signals from the retrieval results.
Approach: They propose to optimize a language model for reformulating search queries in line with retrievers’ preferences by combining a large-scale dataset with Retrievers’ Feedback.
Outcome: The proposed framework outperforms existing methods on two benchmarks and surpasses the state-of-the-art methods.
Optimizing Test-Time Query Representations for Dense Retrieval (2023.findings-acl)

Copied to clipboard

Challenge: Recent developments of dense retrieval rely on quality representations of queries and contexts from pre-trained query and context encoders.
Approach: They propose a test-time optimization of query representations that provides fine-grained pseudo labels over retrieval results.
Outcome: The proposed algorithm improves open-domain question answering accuracy and direct re-ranking by up to 2.0% while running 1.3–2.4x faster with an efficient implementation.
Biomedical Entity Representations with Synonym Marginalization (2020.acl-main)

Copied to clipboard

Challenge: Biomedical named entities are often used as key features in biomedical text mining.
Approach: They propose to use a model-based candidate selection to maximize the marginal likelihood of the synonyms present in top candidates.
Outcome: The proposed model outperforms previous state-of-the-art models on four biomedical entity normalization datasets with three different entity types.
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information (2025.acl-long)

Copied to clipboard

Challenge: Temporal Heads are attention heads that primarily handle temporal knowledge.
Approach: They discover Temporal Heads, specific attention heads that primarily handle temporal knowledge, through circuit analysis.
Outcome: The proposed models can handle temporal knowledge without compromising time-invariant and question-answering performances.
Look at the First Sentence: Position Bias in Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Extractive question answering models are trained to predict start and end positions of answers . recent QA models outperform humans in some datasets due to their simplicity and effectiveness.
Approach: They propose to use prior distribution of answer positions as a bias model to reduce position bias.
Outcome: The proposed model outperforms BERT from 37.48% to 81.64% when trained on a biased SQUAD dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations