Papers by Anne Lauscher
Copied to clipboard
| Challenge: | Existing studies show that incorporating demographic factors in language representations improves performance on downstream NLP tasks. |
| Approach: | They use continuous language modeling and dynamic multi-task learning to adapt pre-trained Transformers to incorporate demographic information into their representations. |
| Outcome: | The proposed model shows that the results are consistent with previous studies. |
Copied to clipboard
| Challenge: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Approach: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Outcome: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
Copied to clipboard
| Challenge: | Initial studies have pointed to the potential for harm due to predictive bias, reflecting and potentially reinforcing cultural stereotypes. |
| Approach: | They conduct a survey among non-cisgender individuals and interviews to establish which harms affected individuals anticipate, and how they would like to be represented. |
| Outcome: | The results show that certain non-cisgender identities are consistently (mis)represented as less human, more stereotyped and more sexualised. |
Copied to clipboard
| Challenge: | citation context analysis (CCA) studies the role and purpose of citations in scientific discourse. |
| Approach: | They construct a first comprehensive context definition based on semantic properties of citing text . they use fine-grained semantic properties to evaluate the definition . |
| Outcome: | The proposed definition shows improvements of up to 25% over state-of-the-art methods. |
Copied to clipboard
| Challenge: | a societal movement towards using gender-fair language exists, but gender-free German is barely supported in machine translation. |
| Approach: | They propose to use a community-created gender-fair language dictionary to study gender-neutral German . they also use encyclopedic text and parliamentary speeches to translate the words in isolation . |
| Outcome: | The proposed study shows that most systems produce mainly masculine forms and rarely gender-neutral variants. |
Copied to clipboard
| Challenge: | Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, but current research focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind. |
| Approach: | They propose a method to mitigate gender bias in machine translation by using a corpus of machine translations from the WinoMT corpus. |
| Outcome: | The proposed model can solve multiple NLP tasks when prompted, but it lacks fairness and ethical considerations. |
Copied to clipboard
| Challenge: | Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each other’s work. |
| Approach: | They propose to use a dataset of 12.6K citation contexts from 1.2K computational linguistics papers to model three important CCA phenomena. |
| Outcome: | The proposed dataset contains 12.6K citation contexts from 1.2K computational linguistics papers and can model these phenomena. |
Copied to clipboard
| Challenge: | Large Language Models struggle to detect lazy thinking in a zero-shot setting, but instruction-based fine-tuning significantly boosts performance by 10-20 performance points. |
| Approach: | They propose to use LazyReview to train junior reviewers in the community to detect lazy thinking in peer-review sentences annotated with fine-grained lazy thinking categories. |
| Outcome: | The proposed dataset shows that LLMs struggle to detect lazy thinking instances in a zero-shot setting, while instruction-based fine-tuning significantly boosts performance by 10-20 performance points. |
Copied to clipboard
| Challenge: | Existing debiasing methods modify all of the PLM parameters, which is costly and leads to (catastrophic) forgetting of useful language knowledge. |
| Approach: | They propose a modular debiasing approach based on dedicated adapters that inject adapter modules into the original PLM layers and update only the adapters. |
| Outcome: | The proposed approach is based on dedicated adapters and retains fairness even after large-scale training. |
Copied to clipboard
| Challenge: | a survey shows that laypeople express different ethical concerns than professionals . acl-code-ethics provides a taxonomy for ethical concerns . |
| Approach: | They propose to annotate a corpus of ethical concern statements from scientific papers . they extract ethical concern keywords from the statements and automate the process . |
| Outcome: | The proposed corpus of ethical concern statements compares with existing taxonomies and guidelines pointing to gaps and actionable insights. |
Copied to clipboard
| Challenge: | Argument quality assessment is critical for opinion formation, decision making, writing education, and the like. |
| Approach: | They propose to use large language models to leverage knowledge across contexts to enable a much more reliable assessment. |
| Outcome: | The proposed approach improves the quality of argumentation and the ability to leverage knowledge across contexts. |
Copied to clipboard
| Challenge: | Existing research on argumentation models does not provide a systematic overview of the types of knowledge required in CA tasks. |
| Approach: | They propose a taxonomy of the types of knowledge required in CA tasks . authors propose exploitation of these knowledge types for four main research areas . |
| Outcome: | The proposed taxonomy proposes a systematic overview of the types of knowledge required in CA tasks. |
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only. |
| Approach: | They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. |
| Outcome: | The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions. |
Copied to clipboard
| Challenge: | despite LLMs becoming increasingly multilingual, most studies on detecting and quantifying LLM hallucination are English-centric . |
| Approach: | They train a multilingual hallucination detection model and conduct a large-scale study across 30 languages and 6 open-source LLM families. |
| Outcome: | The proposed model is based on an English-centric model and annotates gold data for five high-resource languages. |
Copied to clipboard
| Challenge: | a global majority of non-English speakers are underrepresented by large language models . however, most open LLMs are limited in their language coverage . |
| Approach: | They propose a silver standard benchmark for basic open-ended question answering with 27.4k test questions across a typologically diverse set of 137 languages. |
| Outcome: | The proposed model can answer questions in 27.4k questions across 137 languages. |
Copied to clipboard
| Challenge: | Gender-fair language fosters inclusion by addressing all genders or using neutral forms. |
| Approach: | They present a dataset that provides high-quality reformulations for German text classification . they find substantial label flips, reduced prediction certainty, and altered attention patterns . |
| Outcome: | The proposed dataset provides high-quality reformulations for German text classification . it finds label flips, reduced prediction certainty, and significantly altered attention patterns . |
Copied to clipboard
| Challenge: | Recent work suggests that instead of directly countering surface-level reasoning, one should follow an argumentation style inspired by the Jiu-Jitsu “soft” combat system. |
| Approach: | They propose a task of attitude and theme-guided rebuttal generation for peer reviews to enrich existing discourse structure with attitude roots, attitude themes, and canonical reversals. |
| Outcome: | The proposed task is based on an existing dataset for discourse structure in peer reviews with attitude roots, attitude themes, and canonical rebuttals. |
Copied to clipboard
| Challenge: | Prior research on meta-reviewing has treated this as a summarization problem over review reports . prior research demonstrated that decision-makers can be effectively assisted in such scenarios via dialogue agents. |
| Approach: | They propose to use large-scale large-language models to generate synthetic data for meta-reviewing . they then use these data to train dialogue agents tailored for meta review . |
| Outcome: | The proposed method outperforms *off-the-shelf* dialogue agents in meta-reviewing scenarios. |
Copied to clipboard
| Challenge: | Task-oriented dialog (TOD) is arguably one of the most popular natural language processing (NLP) application areas. |
| Approach: | They propose a multilingual multi-domain TOD dataset that spans four languages . they use a framework for multilingual conversational specialization of pretrained language models . |
| Outcome: | The proposed datasets show that they perform better than existing datasets in English . the proposed framework allows for sample-efficient few-shot transfer for TOD tasks . |
Copied to clipboard
| Challenge: | Recent work on stereotypical biases in semantic spaces is still in its infancy . we present a novel resource for bias measurement specifically tailored to argumentation . |
| Approach: | They propose a resource for bias measurement specifically tailored to argumentation . they use argumentative fine-tuning and debiasing to assess intrinsic bias . |
| Outcome: | The proposed approach is more sustainable and parameter-efficient than full fine-tuning . it can remove bias in general and argumentative language models while improving model performance in downstream tasks. |
Copied to clipboard
| Challenge: | Existing work on argument quality (AQ) focuses on overall quality, but there is no large-scale theory-based corpus and corresponding computational models. |
| Approach: | They propose to use a large-scale English multi-domain argumentative writing corpus annotated with theory-based AQ scores to assess argument quality. |
| Outcome: | The proposed methods improve argument quality in three domains and can be used as strong baselines for future work. |
Copied to clipboard
| Challenge: | Recent work has focused on measuring and mitigating bias in pretrained language models. |
| Approach: | They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model . |
| Outcome: | The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions. |
Copied to clipboard
| Challenge: | Scientific publications are argumentative and often adhere to well-trodden rhetorical patterns and argumentation schemes. |
| Approach: | They investigate the link between scientific publications and rhetorical aspects such as discourse categories or citation contexts by coupling rhetorical classifiers with extraction of argumentative components. |
| Outcome: | The proposed models show significant performance gains for different rhetorical analysis tasks. |
Copied to clipboard
| Challenge: | a lack of research on the interplay between fairness and environmental impact is a problem in natural language processing . fairness is prone to encode and amplify stereotypical social biases, according to several studies . |
| Approach: | They evaluate a technique to reduce energy consumption of English NLP models by knowledge distillation for its impact on fairness. |
| Outcome: | The proposed method reduces energy consumption and environmental impact of English NLP models. |
Copied to clipboard
| Challenge: | Pre-trained language models have outperformed other models on a wide range of tasks . however, there is still little understanding of their knowledge of higher-level aspects of language . |
| Approach: | They investigate whether pre-trained language models have knowledge of sociodemographics . they use traditional probing techniques to probe the knowledge of single-GPU PLMs based on multiple English data sets . |
| Outcome: | The results show that pre-trained language models outperform other models on a wide range of tasks. |
Copied to clipboard
| Challenge: | a new study shows that cultural background significantly affects multimodal hate speech moderation models . a limited dataset excludes multi-modal forms of hate and excludes non-English-speaking cultures . the lowest pairwise label agreement between the USA and India is due to cultural factors . |
| Approach: | They use a multimodal and multilingual parallel hate speech dataset to examine cultural differences . they find that cultural background significantly affects multimodal hate speech annotation . |
| Outcome: | The proposed dataset shows that cultural background significantly affects multimodal hate speech annotation. |
Copied to clipboard
| Challenge: | In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes . |
| Approach: | They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias. |
| Outcome: | The proposed model reproduces dialect usage bias in association task and decision task. |
Copied to clipboard
| Challenge: | Existing studies show that multilingual transformers are less effective in resource-lean scenarios and for distant languages. |
| Approach: | They propose to use massively multilingual transformers to pretrain languages . they show that MMTs are less effective in resource-lean scenarios and distant languages if they are pre-trained via language modeling . |
| Outcome: | The proposed model is less effective in resource-lean scenarios and for distant languages than cross-lingual word embeddings. |
Copied to clipboard
| Challenge: | Language models (LMs) may produce toxic text that contains hate speech, insults, or vulgarity, even when prompted with innocuous text. |
| Approach: | They propose an interpretability framework that aligns the behavior of language models based on their outputs and internal representations. |
| Outcome: | The proposed framework bridges behavioral and internal perspectives for toxicity for the first time. |
Copied to clipboard
| Challenge: | Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals. |
| Approach: | They compare 3rd-person pronoun translations to five other languages . they propose to address gender exclusivity in future research . |
| Outcome: | The proposed method compares translations of gendered vs. gender-neutral pronouns from english to five other languages and vice versa. |
Copied to clipboard
| Challenge: | Current modeling of 3rd person pronouns ignores neopronoun phenomena like naive pronounes, which are not (yet) widely established. |
| Approach: | They propose to validate existing and novel approaches for modeling 3rd person pronouns in language technology and validate them through a survey. |
| Outcome: | The proposed model excludes non-binary users, while ignoring gender-specific phenomena. |
Copied to clipboard
| Challenge: | Unsupervised pretraining models encode only distributional knowledge encoded in text corpora, incorporated through language modeling objectives. |
| Approach: | They generalize a standard BERT model to a multi-task learning setting and integrate discrete knowledge on word-level semantic similarity into pretraining. |
| Outcome: | The proposed model outperforms the lexically blind “vanilla” model on several language understanding tasks. |
Copied to clipboard
| Challenge: | Existing studies on temporal knowledge in text-to-image models have not explored how temporal phenomena are handled in text models. |
| Approach: | They propose a data set to holistically evaluate temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
| Outcome: | The proposed model evaluates temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
Copied to clipboard
| Challenge: | Prior work has shown that text-to-image models produce culturally stereotypical depictions when faced with languages other than English . |
| Approach: | They propose to use a set of prompts translated into 14 languages to prompt seven T2I models. |
| Outcome: | The proposed model is compared with seven models in 171 cultural identities translated into 14 languages and shows that all but one model exhibit strong surface-level tendency in at least two languages. |
Copied to clipboard
| Challenge: | Including an image in a multimodal query significantly increases the likelihood of the model returning an English response regardless of the language of the query. |
| Approach: | They propose a two-pronged approach that combines extensive ablation of the design space with a mechanistic analysis of the models’ internal representations of image and text inputs. |
| Outcome: | The proposed approach reduces the multilingual error by switching the language backbone for a bilingual language model. |
Copied to clipboard
| Challenge: | Existing studies have tested language models' ability to reason over time and space in isolation or only in simple or artificial environments. |
| Approach: | They present a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones to evaluate their ability to jointly reason over time and space. |
| Outcome: | The proposed models perform well on reasoning tasks involving only temporal knowledge, but performance remains constrained on tasks that require connecting temporal and geographic information. |
Copied to clipboard
| Challenge: | Existing MT models are limited in size and often consist of single sentences or single gender-fair formulation types. |
| Approach: | They propose a benchmark for machine translation that features extended passages with professional translations implementing gender-fair alternatives: neutral rewording, typographical solutions and neologistic forms. |
| Outcome: | The proposed benchmark features extended passages with professional translations implementing three gender-fair alternatives: neutral rewording, typographical solutions (gender star), and neologistic forms (-ens forms). |
Copied to clipboard
| Challenge: | ACL removed the anonymity period for conference submissions in February 2024 . |
| Approach: | They track preprinting trends for 47k publications and analyze 1.9k peer reviews . they suggest improving visibility and investing in diversity initiatives . |
| Outcome: | The proposed anonymity period was removed in 2024, but it was ineffective for underrepresented researchers . the authors suggest addressing D&I issues rather than implementing anonymity policies. |
Copied to clipboard
| Challenge: | Stereotypical bias encoded in language models (LMs) poses a threat to safe language technology . current research lacks a thorough understanding of manifestations of biases in specific model weights. |
| Approach: | They propose a method that localizes and edits weights associated with gender bias . they use local contrastive editing to localize and control a small subset of weights . |
| Outcome: | The proposed method localizes and controls a small subset of weights that encode gender bias. |
Copied to clipboard
| Challenge: | Recent research has shown that distributional word vector spaces often encode stereotypical human biases, such as racism and sexism. |
| Approach: | They propose a platform that measures and mitigates bias in word embeddings by executing two (mutually composable) debiasing models. |
| Outcome: | The proposed platform can measure and mitiga bias in word embeddings. |
Copied to clipboard
| Challenge: | Multi-task learning (MTL) has shown significant practical benefits when using language models . current two stage MTL introduces a substantial number of additional parameters . |
| Approach: | They propose a multi-task learning method that leverages existing knowledge for a target task. |
| Outcome: | The proposed method outperforms baselines on three benchmarks and two encoder LMs with a small number of transfer parameters. |
Copied to clipboard
| Challenge: | Genderneutral translation (GNT) is a linguistic strategy towards fairer communication across languages. |
| Approach: | They propose to use a multilingual evaluation resource to evaluate inclusive translation with state-of-the-art instruction-following language models (LMs) |
| Outcome: | The proposed model can recognize when neutrality is appropriate, but cannot consistently produce neutral translations, limiting their usability. |
Copied to clipboard
| Challenge: | Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. |
| Approach: | They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data . |
| Outcome: | The proposed model outperforms existing models in 14 tasks and 56 languages. |
Copied to clipboard
| Challenge: | Existing and new approaches to zero-shot knowledge composition are lacking in NLP. |
| Approach: | They propose a framework for zero-shot module composition that unifies existing and some novel variations for selecting, weighting, and combining parameter modules under a single unified notion. |
| Outcome: | The proposed framework enables a systematic unification of concepts and enables the first comprehensive benchmarking study of various zero-shot knowledge composition strategies. |
Copied to clipboard
| Challenge: | Recent studies have focused on the ethical aspects of NLP, but little to no discussion of the terminology and theories underpinning those efforts and their implications. |
| Approach: | They propose to provide an overview of some important ethical concepts stemming from philosophy and to survey the existing literature on moral NLP w.r.t. their findings show that most papers neither provide a clear definition of the terms they use nor adhere to definitions from philosophy. |
| Outcome: | The findings show that most papers neither provide a clear definition of the terms they use nor adhere to definitions from philosophy. |
Copied to clipboard
| Challenge: | Existing studies on sociodemographic prompting have not explored the effectiveness of this technique. |
| Approach: | They propose to use sociodemographic prompting to steer models towards answers that humans with specific sociodemography would give. |
| Outcome: | The proposed technique can improve zero-shot learning by focusing on human sociodemographic profiles. |
Copied to clipboard
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |
Copied to clipboard
| Challenge: | Recent work shows that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining. |
| Approach: | They propose a resource-efficient and modular domain specialization by means of domain adapters in which domain knowledge is encoded. |
| Outcome: | The proposed framework extracts domain-specific terms and then uses them to build DomainCC and DomainReddit resources based on masked language modeling and response selection objectives. |