Papers by Katerina Marazopoulou
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks (2024.acl-long)
Copied to clipboard
| Challenge: | Using benchmarks to evaluate Large Language Models is inconsistent with the assumption that the test prompts within a benchmark represent a random sample from some real-world distribution of interest. |
| Approach: | They propose to use a model's average performance across the test prompts of a benchmark to evaluate its performance. |
| Outcome: | The results show that the correlation between model performance across test prompts and the test prompt can change model rankings on major benchmarks. |