Challenge: grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning.
Approach: a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german .
Outcome: a corpus of German gender-inclusive language is assembled and will be included in the release.

Similar Papers

taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)

Copied to clipboard

Challenge: a large corpus of German newspaper articles is available for free in other languages, such as English.
Approach: They propose to use taz2024full to analyse gender representation across four decades of reporting.
Outcome: The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Gender-fair language fosters inclusion by addressing all genders or using neutral forms.
Approach: They present a dataset that provides high-quality reformulations for German text classification . they find substantial label flips, reduced prediction certainty, and altered attention patterns .
Outcome: The proposed dataset provides high-quality reformulations for German text classification . it finds label flips, reduced prediction certainty, and significantly altered attention patterns .
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German (2024.findings-acl)

Copied to clipboard

Challenge: a societal movement towards using gender-fair language exists, but gender-free German is barely supported in machine translation.
Approach: They propose to use a community-created gender-fair language dictionary to study gender-neutral German . they also use encyclopedic text and parliamentary speeches to translate the words in isolation .
Outcome: The proposed study shows that most systems produce mainly masculine forms and rarely gender-neutral variants.
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)

Copied to clipboard

Challenge: Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases.
Approach: They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French.
Outcome: The proposed system detects gendered forms and replaces them with neutral or opposite forms.
Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing MT models are limited in size and often consist of single sentences or single gender-fair formulation types.
Approach: They propose a benchmark for machine translation that features extended passages with professional translations implementing gender-fair alternatives: neutral rewording, typographical solutions and neologistic forms.
Outcome: The proposed benchmark features extended passages with professional translations implementing three gender-fair alternatives: neutral rewording, typographical solutions (gender star), and neologistic forms (-ens forms).
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)

Copied to clipboard

Challenge: a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women.
Approach: They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations .
Outcome: The proposed method compares different technologies on two languages, English and French.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Identifying and Reducing Gender Bias in Word-Level Language Models (N19-3)

Copied to clipboard

Challenge: Existing discriminatory biases in training data can be amplified by models . text corpora exhibit socially problematic biase .
Approach: They propose a metric to measure gender bias and a regularization loss term to minimize embeddings onto an embeddable subspace that encodes gender.
Outcome: The proposed method reduces gender bias up to an optimal weight assigned to the loss term, and the model becomes unstable as the perplexity increases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations