Papers by Carolin Holtermann
GIMMICK: Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only. |
| Approach: | They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. |
| Outcome: | The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions. |
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ (2024.findings-acl)
Copied to clipboard
| Challenge: | a global majority of non-English speakers are underrepresented by large language models . however, most open LLMs are limited in their language coverage . |
| Approach: | They propose a silver standard benchmark for basic open-ended question answering with 27.4k test questions across a typologically diverse set of 137 languages. |
| Outcome: | The proposed model can answer questions in 27.4k questions across 137 languages. |
Fair and Argumentative Language Modeling for Computational Argumentation (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work on stereotypical biases in semantic spaces is still in its infancy . we present a novel resource for bias measurement specifically tailored to argumentation . |
| Approach: | They propose a resource for bias measurement specifically tailored to argumentation . they use argumentative fine-tuning and debiasing to assess intrinsic bias . |
| Outcome: | The proposed approach is more sustainable and parameter-efficient than full fine-tuning . it can remove bias in general and argumentative language models while improving model performance in downstream tasks. |
Large Language Models Discriminate Against Speakers of German Dialects (2025.emnlp-main)
Copied to clipboard
| Challenge: | In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes . |
| Approach: | They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias. |
| Outcome: | The proposed model reproduces dialect usage bias in association task and decision task. |
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing studies on temporal knowledge in text-to-image models have not explored how temporal phenomena are handled in text models. |
| Approach: | They propose a data set to holistically evaluate temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
| Outcome: | The proposed model evaluates temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation (2026.eacl-long)
Copied to clipboard
| Challenge: | Prior work has shown that text-to-image models produce culturally stereotypical depictions when faced with languages other than English . |
| Approach: | They propose to use a set of prompts translated into 14 languages to prompt seven T2I models. |
| Outcome: | The proposed model is compared with seven models in 171 cultural identities translated into 14 languages and shows that all but one model exhibit strong surface-level tendency in at least two languages. |
Why do LLaVA Vision-Language Models Reply to Images in English? (2024.findings-emnlp)
Copied to clipboard
Musashi Hinck, Carolin Holtermann, Matthew Olson, Florian Schneider, Sungduk Yu, Anahita Bhiwandiwalla, Anne Lauscher, Shao-Yen Tseng, Vasudev Lal
| Challenge: | Including an image in a multimodal query significantly increases the likelihood of the model returning an English response regardless of the language of the query. |
| Approach: | They propose a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models’ internal representations of image and text inputs. |
| Outcome: | The proposed approach reduces the multilingual error by switching the language backbone for a bilingual language model. |
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have tested language models' ability to reason over time and space in isolation or only in simple or artificial environments. |
| Approach: | They present a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones to evaluate their ability to jointly reason over time and space. |
| Outcome: | The proposed models perform well on reasoning tasks involving only temporal knowledge, but performance remains constrained on tasks that require connecting temporal and geographic information. |
ScaLearn: Simple and Highly Parameter-Efficient Task Transfer by Learning to Scale (2024.findings-acl)
Copied to clipboard
| Challenge: | Multi-task learning (MTL) has shown significant practical benefits when using language models . current two stage MTL introduces a substantial number of additional parameters . |
| Approach: | They propose a multi-task learning method that leverages existing knowledge for a target task. |
| Outcome: | The proposed method outperforms baselines on three benchmarks and two encoder LMs with a small number of transfer parameters. |
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)
Copied to clipboard
Gregor Geigle, Florian Schneider, Carolin Holtermann, Chris Biemann, Radu Timofte, Anne Lauscher, Goran Glavaš
| Challenge: | Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. |
| Approach: | They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data . |
| Outcome: | The proposed model outperforms existing models in 14 tasks and 56 languages. |
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing and new approaches to zero-shot knowledge composition are lacking in NLP. |
| Approach: | They propose a framework for zero-shot module composition that unifies existing and some novel variations for selecting, weighting, and combining parameter modules under a single unified notion. |
| Outcome: | The proposed framework enables a systematic unification of concepts and enables the first comprehensive benchmarking study of various zero-shot knowledge composition strategies. |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |