Papers by Omer Akgul
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, resulting in inconsistent or unreliable generated text. |
| Approach: | They propose a logit-based ensemble method to measure LLM consistency and propose to use it to evaluate human ratings of LLM reliability. |
| Outcome: | The proposed method matches the best-performing existing metric in estimating human ratings of LLM consistency. |