Papers by Suyoung Bae
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets. |
| Approach: | They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries. |
| Outcome: | The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art. |
EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing RL methods assign query-level rewards to all clauses, treating correct and incorrect clauses equally. |
| Approach: | They propose a method which provides fine-grained supervision through clause-level rewards. |
| Outcome: | Experiments on widely-used Text-to-SQL benchmarks show that EXPO-SqL outperforms existing methods by fine-grained clause-level learning. |
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing zero-shot methods for Question Answering (QA) are efficient but fail to consider context and prevent bias propagation in the answers. |
| Approach: | They propose a method for debiasing Large Language Models using context-adaptive prompt generation that takes appropriate debiased actions based on the context and aNeutral Answer Guidance Generation to suppress the LLMs make objective judgments about the context. |
| Outcome: | The proposed method achieves state-of-the-art zero-shot debiased QA performance across eight LLMs. |
CharMoral: A Character Morality Dataset for Morally Dynamic Character Analysis in Long-Form Narratives (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies on character analysis focus on character identification, social network analysis, and the exploration of characters' personas or personalities. |
| Approach: | They propose a four-stage framework to automatically classify actions as moral or immoral based on context. |
| Outcome: | The proposed framework is effective in moral reasoning tasks in multiple genres. |
UCGRec: User-Centric Graph Learning for LLM-based Sequential Recommendation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for sequential recommendation rely primarily on item descriptions or utilize user preferences independently. |
| Approach: | They propose a method that integrates diverse user-relevant preference signals into a unified user-centric graph and injects the graph-based knowledge into the LLM through end-to-end training with graph neural networks. |
| Outcome: | The proposed method outperforms conventional and state-of-the-art methods on four widely used sequential real-world recommendation datasets. |
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data (2025.naacl-long)
Copied to clipboard
| Challenge: | In many natural language processing tasks, model training often leads to spurious correlations . shortcuts allow models to rely on irrelevant patterns in the data, leading to biased predictions. |
| Approach: | They propose a method to generate structure-aware positive and negative sentences using tagging. |
| Outcome: | The proposed method improves model robustness and generalization across different environments while minimizing spurious correlations. |