Papers by Subin An
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing methods to evaluate open-ended survey responses are expensive and lack ground-truth reference for comparison. |
| Approach: | They propose a two-stage evaluation framework specifically designed for human survey responses that uses gibberish filtering to remove nonsensical responses. |
| Outcome: | The proposed evaluation framework outperforms existing metrics on English and Korean datasets and shows strong correlations with expert assessment. |
FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and Korean (2025.emnlp-main)
Copied to clipboard
| Challenge: | Figurative language is a core component of everyday communication . existing benchmarks focus on sentence-level classification or inference tasks . |
| Approach: | They propose a multilingual benchmark that evaluates figurative usage in dialogue . they use a sentence-level diagnostic task to embed figurativ choices into multi-turn contexts . |
| Outcome: | The benchmark evaluates large language models' ability to use figurative expressions coherently in conversation. |