Papers by Aykut Erdem
CRAFT: A Benchmark for Causal Reasoning About Forces and inTeractions (2022.findings-acl)
Copied to clipboard
Tayfun Ates, M. Ateşoğlu, Çağatay Yiğit, Ilker Kesen, Mert Kobas, Erkut Erdem, Aykut Erdem, Tilbe Goksun, Deniz Yuret
| Challenge: | Existing models with similar physical and causal understanding capabilities are still underdeveloped. |
| Approach: | They propose a video question answering dataset that requires causal reasoning about physical forces and object interactions. |
| Outcome: | The proposed dataset requires causal reasoning about physical forces and object interactions. |
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing Turkish benchmarks lack task diversity or culturally relevant content . Cetvel combines a broad range of discriminative and generative tasks . |
| Approach: | They propose a benchmark to evaluate large language models in Turkish . Cetvel combines a broad range of discriminative and generative tasks . they find that Turkish-centric instruction-tuned models generally underperform . |
| Outcome: | The proposed benchmark covers 23 tasks grouped into seven categories . it shows that Turkish-centric instruction-tuned models underperform relative to multilingual or general-purpose models despite being tailored for the language. |
DeVisE: Towards the Behavioral Testing of Medical Large Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluations of large language models do not reveal whether their outputs reflect genuine medical reasoning or superficial correlations. |
| Approach: | They propose a framework that probes fine-grained clinical understanding through controlled counterfactuals. |
| Outcome: | The proposed framework is based on demographic and vital signs data from the ICU discharge notes of patients in the intensive care unit (MIMIC-IV). |
Sequential Compositional Generalization in Multimodal Models (2024.naacl-long)
Copied to clipboard
| Challenge: | a growing number of multimodal models have a limited capacity for generalization . however, prior studies into compositionality have focused on visual grounding and downstream tasks like image captioning. |
| Approach: | They examine compositional generalization using egocentric kitchen activity videos . they find bi-modal and tri-modal models exhibit a clear edge over their text-only counterparts . |
| Outcome: | The proposed model outperforms text-only models in a multimodal setting. |
RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes (D18-1)
Copied to clipboard
| Challenge: | Existing comprehension tests for QA are limited by the text sources and questionanswer formats. |
| Approach: | They propose a dataset for multimodal comprehension of cooking recipes . preliminary results indicate RecipeQA will serve as a challenging test bed . |
| Outcome: | The proposed dataset will serve as a test bed and ideal benchmark for evaluating machine comprehension systems. |
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)
Copied to clipboard
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia
| Challenge: | Pre-trained language models have been shown to improve performance in many natural language tasks. |
| Approach: | They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora. |
| Outcome: | The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation. |
Harnessing Dataset Cartography for Improved Compositional Generalization in Transformers (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to understanding compositional generalization of models have focused on novel architectures and alternative learning paradigms. |
| Approach: | They propose a method that harnesses the power of dataset cartography to improve model accuracy by strategically identifying a subset of compositional generalization data. |
| Outcome: | The proposed method improves model accuracy by 10% on CFQ and COGS datasets. |