Papers by Katerina Marazopoulou

1 papers
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks (2024.acl-long)

Copied to clipboard

Challenge: Using benchmarks to evaluate Large Language Models is inconsistent with the assumption that the test prompts within a benchmark represent a random sample from some real-world distribution of interest.
Approach: They propose to use a model's average performance across the test prompts of a benchmark to evaluate its performance.
Outcome: The results show that the correlation between model performance across test prompts and the test prompt can change model rankings on major benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations