Papers by Alaa Elsetohy
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling (2026.acl-long)
Copied to clipboard
Alaa Elsetohy, Sama Hadhoud, Haryo Akbarianto Wibowo, Chenxi Whitehouse, Genta Indra Winata, Fajri Koto, Alham Fikri Aji
| Challenge: | Existing benchmarks test reasoning over culturally grounded premises, but translation-parallel benchmarks inherit English-centric scenarios. |
| Approach: | They propose a template-first benchmark that factorizes reasoning type and cultural aspect across question languages. |
| Outcome: | The proposed benchmark factorizes reasoning type and cultural aspect across question languages. |
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming (2026.findings-acl)
Copied to clipboard
Sama Hadhoud, Alaa Elsetohy, Frederikus Hudi, Jan Christian Blaise Cruz, Steven Halim, Alham Fikri Aji
| Challenge: | Existing evaluations conflate algorithmic reasoning with code-level implementation. |
| Approach: | They propose to center editorials in both solution generation and evaluation . they propose to compare editorials to gold standards and validate an LLM-as-a-judge protocol . |
| Outcome: | The proposed approach improves solve rates on some LLMs with gold editorials . but the gap between gold and generated editorials shows bottlenecks in implementation . |