Papers by Sacha Muller
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | Existing automated RAG evaluation frameworks overlook important failure modes when using GPT-4 as a judge. |
| Approach: | They propose a novel pipeline to assess the calibration and discrimination capabilities of judge models by using a meta-evaluation benchmark of 144 unit tests to identify key failure modes. |
| Outcome: | The proposed pipeline improves on existing frameworks, while state-of-the-art open-source judges do not generalize to their proposed criteria. |