Papers by Sultan AlRashed

2 papers
Cards Against Contamination: TCG-Bench for Difficulty-Scalable Multilingual LLM Reasoning (2026.findings-eacl)

Copied to clipboard

Challenge: Recent studies find 25-50% of evaluation datasets appear in training corpora . contamination hinders the possibility to differentiate memorization and reasoning skills.
Approach: They propose a two-player trading card game that is contaminated by a public engine and hidden card implementations to prevent benchmark saturation.
Outcome: The proposed benchmark is based on a new two-player trading card game similar to Magic: The Gathering.
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards (2024.acl-long)

Copied to clipboard

Challenge: Existing leaderboards are often taken at face value, but this is costly . a recent study shows that minor perturbations to the benchmark result in rankings up to 8 positions.
Approach: They propose to use a *hybrid* scoring method for answer selection for large language models . they find that minor perturbations to the benchmark result in rankings changes .
Outcome: The proposed model is a hybrid scoring method, the authors argue . the proposed model could be used to improve the performance of large language models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations