Papers by Floris Weers
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? (2025.acl-long)
Copied to clipboard
| Challenge: | Pairwise feedback is widely used to evaluate and provide feedback to large language models (LLMs). |
| Approach: | They propose a tool-using agentic system to provide higher quality feedback on three challenging response domains: long-form factual, math and code tasks. |
| Outcome: | The proposed system can provide higher quality pairwise comparisons on three domains, independent of the LLM’s internal knowledge and biases. |