Papers by Alaa Elsetohy

2 papers
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks test reasoning over culturally grounded premises, but translation-parallel benchmarks inherit English-centric scenarios.
Approach: They propose a template-first benchmark that factorizes reasoning type and cultural aspect across question languages.
Outcome: The proposed benchmark factorizes reasoning type and cultural aspect across question languages.
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations conflate algorithmic reasoning with code-level implementation.
Approach: They propose to center editorials in both solution generation and evaluation . they propose to compare editorials to gold standards and validate an LLM-as-a-judge protocol .
Outcome: The proposed approach improves solve rates on some LLMs with gold editorials . but the gap between gold and generated editorials shows bottlenecks in implementation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations