Papers by Carolin Holtermann

12 papers
GIMMICK: Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only.
Approach: They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions.
Outcome: The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions.
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ (2024.findings-acl)

Copied to clipboard

Challenge: a global majority of non-English speakers are underrepresented by large language models . however, most open LLMs are limited in their language coverage .
Approach: They propose a silver standard benchmark for basic open-ended question answering with 27.4k test questions across a typologically diverse set of 137 languages.
Outcome: The proposed model can answer questions in 27.4k questions across 137 languages.
Fair and Argumentative Language Modeling for Computational Argumentation (2022.acl-long)

Copied to clipboard

Challenge: Recent work on stereotypical biases in semantic spaces is still in its infancy . we present a novel resource for bias measurement specifically tailored to argumentation .
Approach: They propose a resource for bias measurement specifically tailored to argumentation . they use argumentative fine-tuning and debiasing to assess intrinsic bias .
Outcome: The proposed approach is more sustainable and parameter-efficient than full fine-tuning . it can remove bias in general and argumentative language models while improving model performance in downstream tasks.
Large Language Models Discriminate Against Speakers of German Dialects (2025.emnlp-main)

Copied to clipboard

Challenge: In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes .
Approach: They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias.
Outcome: The proposed model reproduces dialect usage bias in association task and decision task.
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on temporal knowledge in text-to-image models have not explored how temporal phenomena are handled in text models.
Approach: They propose a data set to holistically evaluate temporal knowledge in image generation using 7.9k prompts and more than 600 reference images.
Outcome: The proposed model evaluates temporal knowledge in image generation using 7.9k prompts and more than 600 reference images.
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation (2026.eacl-long)

Copied to clipboard

Challenge: Prior work has shown that text-to-image models produce culturally stereotypical depictions when faced with languages other than English .
Approach: They propose to use a set of prompts translated into 14 languages to prompt seven T2I models.
Outcome: The proposed model is compared with seven models in 171 cultural identities translated into 14 languages and shows that all but one model exhibit strong surface-level tendency in at least two languages.
Why do LLaVA Vision-Language Models Reply to Images in English? (2024.findings-emnlp)

Copied to clipboard

Challenge: Including an image in a multimodal query significantly increases the likelihood of the model returning an English response regardless of the language of the query.
Approach: They propose a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models’ internal representations of image and text inputs.
Outcome: The proposed approach reduces the multilingual error by switching the language backbone for a bilingual language model.
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have tested language models' ability to reason over time and space in isolation or only in simple or artificial environments.
Approach: They present a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones to evaluate their ability to jointly reason over time and space.
Outcome: The proposed models perform well on reasoning tasks involving only temporal knowledge, but performance remains constrained on tasks that require connecting temporal and geographic information.
ScaLearn: Simple and Highly Parameter-Efficient Task Transfer by Learning to Scale (2024.findings-acl)

Copied to clipboard

Challenge: Multi-task learning (MTL) has shown significant practical benefits when using language models . current two stage MTL introduces a substantial number of additional parameters .
Approach: They propose a multi-task learning method that leverages existing knowledge for a target task.
Outcome: The proposed method outperforms baselines on three benchmarks and two encoder LMs with a small number of transfer parameters.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition (2024.findings-eacl)

Copied to clipboard

Challenge: Existing and new approaches to zero-shot knowledge composition are lacking in NLP.
Approach: They propose a framework for zero-shot module composition that unifies existing and some novel variations for selecting, weighting, and combining parameter modules under a single unified notion.
Outcome: The proposed framework enables a systematic unification of concepts and enables the first comprehensive benchmarking study of various zero-shot knowledge composition strategies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations