Papers by Aykut Erdem

7 papers
CRAFT: A Benchmark for Causal Reasoning About Forces and inTeractions (2022.findings-acl)

Copied to clipboard

Challenge: Existing models with similar physical and causal understanding capabilities are still underdeveloped.
Approach: They propose a video question answering dataset that requires causal reasoning about physical forces and object interactions.
Outcome: The proposed dataset requires causal reasoning about physical forces and object interactions.
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish (2026.eacl-long)

Copied to clipboard

Challenge: Existing Turkish benchmarks lack task diversity or culturally relevant content . Cetvel combines a broad range of discriminative and generative tasks .
Approach: They propose a benchmark to evaluate large language models in Turkish . Cetvel combines a broad range of discriminative and generative tasks . they find that Turkish-centric instruction-tuned models generally underperform .
Outcome: The proposed benchmark covers 23 tasks grouped into seven categories . it shows that Turkish-centric instruction-tuned models underperform relative to multilingual or general-purpose models despite being tailored for the language.
DeVisE: Towards the Behavioral Testing of Medical Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing evaluations of large language models do not reveal whether their outputs reflect genuine medical reasoning or superficial correlations.
Approach: They propose a framework that probes fine-grained clinical understanding through controlled counterfactuals.
Outcome: The proposed framework is based on demographic and vital signs data from the ICU discharge notes of patients in the intensive care unit (MIMIC-IV).
Sequential Compositional Generalization in Multimodal Models (2024.naacl-long)

Copied to clipboard

Challenge: a growing number of multimodal models have a limited capacity for generalization . however, prior studies into compositionality have focused on visual grounding and downstream tasks like image captioning.
Approach: They examine compositional generalization using egocentric kitchen activity videos . they find bi-modal and tri-modal models exhibit a clear edge over their text-only counterparts .
Outcome: The proposed model outperforms text-only models in a multimodal setting.
RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes (D18-1)

Copied to clipboard

Challenge: Existing comprehension tests for QA are limited by the text sources and questionanswer formats.
Approach: They propose a dataset for multimodal comprehension of cooking recipes . preliminary results indicate RecipeQA will serve as a challenging test bed .
Outcome: The proposed dataset will serve as a test bed and ideal benchmark for evaluating machine comprehension systems.
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models have been shown to improve performance in many natural language tasks.
Approach: They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora.
Outcome: The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation.
Harnessing Dataset Cartography for Improved Compositional Generalization in Transformers (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to understanding compositional generalization of models have focused on novel architectures and alternative learning paradigms.
Approach: They propose a method that harnesses the power of dataset cartography to improve model accuracy by strategically identifying a subset of compositional generalization data.
Outcome: The proposed method improves model accuracy by 10% on CFQ and COGS datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations