Papers by Katharina Hämmerl

6 papers
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image (T2I) generation models have great results in image quality, flexibility, and text alignment, but they suffer from substantial gender bias.
Approach: They propose a benchmark to study gender bias in multilingual T2I models . they use multilingual prompts to account for grammatical differences influencing gender .
Outcome: The proposed benchmark shows strong gender biases and language-specific differences across models.
Beyond Literal Token Overlap: Token Alignability for Multilinguality (2025.naacl-short)

Copied to clipboard

Challenge: Existing studies have shown that token overlap is a strong predictor of multilinguality and cross-lingual knowledge transfer between languages with different scripts.
Approach: They propose a subword token alignability metric to understand the impact and quality of multilingual tokenisation.
Outcome: The proposed metric predicts multilinguality much better when scripts are disparate and the overlap of literal tokens is low.
Combining Static and Contextualised Multilingual Embeddings (2022.findings-acl)

Copied to clipboard

Challenge: Static embeddings are less expressive than contextual language models, but can be more straightforwardly aligned across multiple languages.
Approach: They extract static embeddings for 40 languages from XLM-R and validate them with cross-lingual word retrieval and then align them using VecMap.
Outcome: The proposed approach improves multilingual representations by leveraging static embeddings and a pre-training code.
Improving Parallel Sentence Mining for Low-Resource and Endangered Languages (2025.acl-short)

Copied to clipboard

Challenge: Parallel sentence mining is a technique used to find matching sentence pairs from a source and target language.
Approach: They propose a benchmark dataset for parallel sentence mining on three low-resource languages . they apply alignment post-processing and cluster-based isotropy enhancement techniques to one of them .
Outcome: The proposed datasets show better mining quality overall for low-resource languages . the proposed methods are crucial for optimizing parallel data extraction for low resource languages - a new study shows.
A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using Prompts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to access linguistic information in pre-trained multilingual language models are difficult to use.
Approach: They propose prompting and formulate linguistic tasks to test the LM's access to explicit grammar principles and find out what type of information can be obtained .
Outcome: The proposed method can provide access to linguistic features in pre-trained models, but some are harder to capture .
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations