Papers by Sacha Muller

1 papers
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing automated RAG evaluation frameworks overlook important failure modes when using GPT-4 as a judge.
Approach: They propose a novel pipeline to assess the calibration and discrimination capabilities of judge models by using a meta-evaluation benchmark of 144 unit tests to identify key failure modes.
Outcome: The proposed pipeline improves on existing frameworks, while state-of-the-art open-source judges do not generalize to their proposed criteria.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations