Papers with DeepEval
InstaJudge: Aligning Judgment Bias of LLM-as-Judge with Humans in Industry Applications (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Automated evaluation using LLM-as-Judge is a viable alternative to human evaluation, but misalignment of judgment biases between humans and LLMs hinders its use in real-world applications. |
| Approach: | They propose an LLM-as-Judge library that improves alignments of judgment biases through automatic prompt optimization. |
| Outcome: | The proposed library outperforms existing LLM-as-Judge libraries by a large margin while being more cost efficient. |