Papers by Yousef Almushayqih
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards (2024.acl-long)
Copied to clipboard
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, Haidar Khan
| Challenge: | Existing leaderboards are often taken at face value, but this is costly . a recent study shows that minor perturbations to the benchmark result in rankings up to 8 positions. |
| Approach: | They propose to use a *hybrid* scoring method for answer selection for large language models . they find that minor perturbations to the benchmark result in rankings changes . |
| Outcome: | The proposed model is a hybrid scoring method, the authors argue . the proposed model could be used to improve the performance of large language models . |
AraEval: An Arabic Multi-Task Evaluation Suite for Large Language Models (2025.emnlp-main)
Copied to clipboard
Alhanoof Althnian, Norah A. Alzahrani, Shaykhah Z. Alsubaie, Eman Albilali, Ahmed Abdelali, Nouf M. Alotaibi, M Saiful Bari, Yazeed Alnumay, Abdulhamed Alothaimen, Maryam Saif, Shahad D. Alzaidi, Faisal Abdulrahman Mirza, Yousef Almushayqih, Mohammed Al Saleem, Ghadah Alabduljabbar, Abdulmohsen Al-Thubaity, Areeb Alowisheq, Nora Al-Twairesh
| Challenge: | AraEval is a suite of evaluation tasks designed to assess the advanced knowledge, reasoning, truthfulness, and instruction following capabilities of large language models. |
| Approach: | They propose to use AraEval to assess the advanced knowledge, reasoning, truthfulness, and instruction following capabilities of large language models in the Arabic context. |
| Outcome: | The evaluation suite covers a broad spectrum of domains, including science, history, religion, and literature. |