Papers by Katharina Hämmerl
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes (2025.acl-long)
Copied to clipboard
Felix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack, Jindřich Libovický, Alexander Fraser, Kristian Kersting
| Challenge: | Text-to-image (T2I) generation models have great results in image quality, flexibility, and text alignment, but they suffer from substantial gender bias. |
| Approach: | They propose a benchmark to study gender bias in multilingual T2I models . they use multilingual prompts to account for grammatical differences influencing gender . |
| Outcome: | The proposed benchmark shows strong gender biases and language-specific differences across models. |
Beyond Literal Token Overlap: Token Alignability for Multilinguality (2025.naacl-short)
Copied to clipboard
| Challenge: | Existing studies have shown that token overlap is a strong predictor of multilinguality and cross-lingual knowledge transfer between languages with different scripts. |
| Approach: | They propose a subword token alignability metric to understand the impact and quality of multilingual tokenisation. |
| Outcome: | The proposed metric predicts multilinguality much better when scripts are disparate and the overlap of literal tokens is low. |
Combining Static and Contextualised Multilingual Embeddings (2022.findings-acl)
Copied to clipboard
| Challenge: | Static embeddings are less expressive than contextual language models, but can be more straightforwardly aligned across multiple languages. |
| Approach: | They extract static embeddings for 40 languages from XLM-R and validate them with cross-lingual word retrieval and then align them using VecMap. |
| Outcome: | The proposed approach improves multilingual representations by leveraging static embeddings and a pre-training code. |
Improving Parallel Sentence Mining for Low-Resource and Endangered Languages (2025.acl-short)
Copied to clipboard
| Challenge: | Parallel sentence mining is a technique used to find matching sentence pairs from a source and target language. |
| Approach: | They propose a benchmark dataset for parallel sentence mining on three low-resource languages . they apply alignment post-processing and cluster-based isotropy enhancement techniques to one of them . |
| Outcome: | The proposed datasets show better mining quality overall for low-resource languages . the proposed methods are crucial for optimizing parallel data extraction for low resource languages - a new study shows. |
A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using Prompts (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to access linguistic information in pre-trained multilingual language models are difficult to use. |
| Approach: | They propose prompting and formulate linguistic tasks to test the LM's access to explicit grammar principles and find out what type of information can be obtained . |
| Outcome: | The proposed method can provide access to linguistic features in pre-trained models, but some are harder to capture . |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |