GIL-GALaD: Gender Inclusive Language - German Auto-Assembled Large Database (2024.lrec-main)
Copied to clipboard
| Challenge: | grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning. |
| Approach: | a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german . |
| Outcome: | a corpus of German gender-inclusive language is assembled and will be included in the release. |
Similar Papers
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)
Copied to clipboard
| Challenge: | a large corpus of German newspaper articles is available for free in other languages, such as English. |
| Approach: | They propose to use taz2024full to analyse gender representation across four decades of reporting. |
| Outcome: | The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification (2024.emnlp-main)
Copied to clipboard
| Challenge: | Gender-fair language fosters inclusion by addressing all genders or using neutral forms. |
| Approach: | They present a dataset that provides high-quality reformulations for German text classification . they find substantial label flips, reduced prediction certainty, and altered attention patterns . |
| Outcome: | The proposed dataset provides high-quality reformulations for German text classification . it finds label flips, reduced prediction certainty, and significantly altered attention patterns . |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German (2024.findings-acl)
Copied to clipboard
| Challenge: | a societal movement towards using gender-fair language exists, but gender-free German is barely supported in machine translation. |
| Approach: | They propose to use a community-created gender-fair language dictionary to study gender-neutral German . they also use encyclopedic text and parliamentary speeches to translate the words in isolation . |
| Outcome: | The proposed study shows that most systems produce mainly masculine forms and rarely gender-neutral variants. |
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)
Copied to clipboard
| Challenge: | Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases. |
| Approach: | They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French. |
| Outcome: | The proposed system detects gendered forms and replaces them with neutral or opposite forms. |
Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing MT models are limited in size and often consist of single sentences or single gender-fair formulation types. |
| Approach: | They propose a benchmark for machine translation that features extended passages with professional translations implementing gender-fair alternatives: neutral rewording, typographical solutions and neologistic forms. |
| Outcome: | The proposed benchmark features extended passages with professional translations implementing three gender-fair alternatives: neutral rewording, typographical solutions (gender star), and neologistic forms (-ens forms). |
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)
Copied to clipboard
| Challenge: | a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women. |
| Approach: | They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations . |
| Outcome: | The proposed method compares different technologies on two languages, English and French. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
Identifying and Reducing Gender Bias in Word-Level Language Models (N19-3)
Copied to clipboard
| Challenge: | Existing discriminatory biases in training data can be amplified by models . text corpora exhibit socially problematic biase . |
| Approach: | They propose a metric to measure gender bias and a regularization loss term to minimize embeddings onto an embeddable subspace that encodes gender. |
| Outcome: | The proposed method reduces gender bias up to an optimal weight assigned to the loss term, and the model becomes unstable as the perplexity increases. |