Papers by Salaheddin Alzubi
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation (2024.emnlp-main)
Copied to clipboard
| Challenge: | evaluating large language models' output is difficult due to the high cost of human evaluation. |
| Approach: | They propose a family of foundational large autorater models that train on over 100 quality assessment tasks. |
| Outcome: | The proposed model outperforms models on 8 of 12 autorater benchmarks on 53 quality assessment tasks. |