ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments (2025.findings-emnlp)
Copied to clipboard
| Challenge: | LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. |
| Approach: | They propose a stochastic method of moments evaluation over the space of meaning-preserving prompt perturbations and propose resamplings to estimate the number of prompt re-sampleds needed to obtain meaningful results. |
| Outcome: | The proposed method is model-, task-, and metric-agnostic, offering a recipe for meaningful and robust evaluation. |
Similar Papers
Revisiting the Reliability of Language Models in Instruction-Following (2026.acl-long)
Copied to clipboard
| Challenge: | Several benchmarks have been proposed to measure instruction-following accuracy, but these scores do not translate to reliable services in real-world use. |
| Approach: | They propose a new metric reliable@k and develop an automated pipeline to generate cousin prompts. |
| Outcome: | The proposed model can be instantiated with cousin prompts and generates high-quality cousin prompt data. |
FactEval: Evaluating the Robustness of Fact Verification Systems in the Era of Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have made significant advances in every natural language processing task, but they are vulnerable to small perturbations in the inputs, raising concerns about their robustness in the real world. |
| Approach: | They propose a large-scale benchmark for extensive evaluation of LLMs in the fact verification domain covering 17 realistic word-level and character-level perturbations and 4 types of subpopulations. |
| Outcome: | The proposed model is brittle to small input changes and exhibits performance variations across different subpopulations. |
All Prompts Are Created Equal? Evaluating Robustness of LLM Judges Against Non-Adversarial Prompt Variations (2026.findings-acl)
Copied to clipboard
| Challenge: | Our work establishes new standards for assessing LLM judge reliability beyond simple accuracy metrics. |
| Approach: | They propose a diagnostic framework grounded in attribute verifiability that enables principled decisions about evaluation automation. |
| Outcome: | The proposed framework establishes new standards for assessing LLM judge reliability beyond simple accuracy metrics. |
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are sensitive to subtle, non-semantic variations in prompt phrasing and formatting. |
| Approach: | They propose to evaluate 4 methods for improving prompt robustness within a unified experimental framework. |
| Outcome: | The proposed methods are compared to 8 models from Llama, Qwen and Gemma families and are generalized against multiple types of distribution shifts. |
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts (2025.coling-main)
Copied to clipboard
| Challenge: | Existing frameworks for evaluating robustness of large language models rely on standardized benchmarks that can escalate costs and limit evaluations across domains. |
| Approach: | They propose a framework to evaluate the robustness of large language models using adversarial prompts and domain-constrained knowledge guidelines. |
| Outcome: | The proposed framework reduces dependency on conventional benchmarks and provides efficient evaluations in constrained domains. |
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent research has neglected instances-level prompt variations and their implications on subjective evaluations. |
| Approach: | They propose a framework to evaluate and comprehend prompt sensitivity in large language models. |
| Outcome: | The proposed framework evaluates and comprehends prompt sensitivity in large language models. |
Noisy Exemplars Make Large Language Models More Robust: A Domain-Agnostic Behavioral Analysis (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the robustness of LLMs with few-shot prompting techniques are limited. |
| Approach: | They propose to test the robustness of LLMs in multi-hop reasoning tasks via domain-agnostic perturbations. |
| Outcome: | The proposed model is more sensitive to certain perturbations such as replacing words with synonyms and more robust to few-shot prompting methods. |
POSIX: A Prompt Sensitivity Index For Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are sensitive to minor variations in prompts, such as spelling errors, alteration of wording or the prompt template. |
| Approach: | They propose a PrOmpt Sensitivity IndeX to measure prompt sensitivity . they use this to compare prompt sensitability of various open source LLMs . |
| Outcome: | The proposed method can measure and compare prompt sensitivity of open source LLMs. |
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models (2025.naacl-industry)
Copied to clipboard
| Challenge: | Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup. |
| Approach: | They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups. |
| Outcome: | The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency. |
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | a high prompt sensitivity has been widely accepted as a core limitation of large language models . a recent study suggests that prompt senescence may be an artifact of evaluation processes . |
| Approach: | They examine whether prompt sensitivity is an inherent weakness or an artifact of evaluation . they find that heuristic evaluation methods overlook semantically correct responses . large language models have achieved remarkable success across a wide range of tasks . |
| Outcome: | The proposed model is more robust to prompt templates than previously thought . the authors show that prompt sensitivity may be an artifact of evaluation rather than a flaw . |