Papers by Sebastian Möller
Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)
Copied to clipboard
| Challenge: | Evaluating translation models is a trade-off between effort and detail. |
| Approach: | They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts. |
| Outcome: | The proposed method exposes systematic differences between human and machine translations to human experts. |
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem (2025.coling-main)
Copied to clipboard
| Challenge: | Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions. |
| Approach: | They propose a role-modeling approach that employs two LLMs as generator and critic to generate and refine NLEs. |
| Outcome: | The proposed model outperforms self-refine and can perform with less powerful LLMs. |
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing systems based on large language models (LLMs) are more precise and reliable in identifying users’ intentions, but the recognition of intents still presents a challenge in the case of ConvXAI, since little training data exist and the domain is highly specific. |
| Approach: | They propose to use a dataset in the NLP domain for user intent recognition in ConvXAI to improve parsing performance. |
| Outcome: | The proposed system outperforms existing methods and improves on existing ones. |
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation (2026.acl-long)
Copied to clipboard
Qianli Wang, Van Bach Nguyen, Yihong Liu, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schuetze, Sebastian Möller, Vera Schmitt
| Challenge: | Large language models excel at generating English counterfactuals but their effectiveness in generating multilingual counterfacts remains unclear. |
| Approach: | They conduct automatic evaluations on both directly generated and derived counterfactuals in six languages and find that cross-lingual perturbations follow common strategic principles. |
| Outcome: | The proposed models show that translation-based counterfactuals offer higher validity than their directly generated counterparts, but still fall short of matching the quality of the original English counterf actuals. |
From Witch’s Shot to Music Making Bones - Resources for Medical Laymen to Technical Language and Vice Versa (2020.lrec-1)
Copied to clipboard
| Challenge: | Information we share online unveils directly or indirectly information about our lifestyle and health situation. |
| Approach: | They propose a dataset which annotates medical laymen and technical expressions in a patient forum and a set of medical synonyms and definitions. |
| Outcome: | The proposed dataset annotates medical laymen and technical expressions in a patient forum along with a set of medical synonyms and definitions. |
An Empirical Comparison of Question Classification Methods for Question Answering Systems (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for Question Classification are monolingual, but they are not suitable for low-resourced languages. |
| Approach: | They propose to classify the most recent methods in four different categories . they propose to use a low, medium, high, and very high level of dependency on external resources . |
| Outcome: | The proposed method outperforms methods not suitable for low-resource languages. |
Subjective Text Complexity Assessment for German (2022.lrec-1)
Copied to clipboard
| Challenge: | Often, readability is defined as how easily a written text is to read. |
| Approach: | They propose to use a corpus of sentences provided by a German IT service provider to assess the readability of German text. |
| Outcome: | The proposed model can predict complexity of German text by using linguistically motivated features. |
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation (2025.findings-acl)
Copied to clipboard
Qianli Wang, Nils Feldhus, Simon Ostermann, Luis Felipe Villa-Arenas, Sebastian Möller, Vera Schmitt
| Challenge: | Existing frameworks for counterfactual examples are lacking for many tasks. |
| Approach: | They propose a faithful approach for leveraging important words from feature attribution methods to generate counterfactual examples in a zero-shot setting. |
| Outcome: | The proposed framework outperforms state-of-the-art frameworks on many tasks. |
A Linguistically Motivated Test Suite to Semi-Automatically Evaluate German–English Machine Translation Output (2022.lrec-1)
Copied to clipboard
Vivien Macketanz, Eleftherios Avramidis, Aljoscha Burchardt, He Wang, Renlong Ai, Shushen Manakhimova, Ursula Strohriegel, Sebastian Möller, Hans Uszkoreit
| Challenge: | Using fine-grained evaluation techniques, translation outputs have become better and more fluent. |
| Approach: | They propose a fine-grained test suite for the language pair German–English . they describe the creation and implementation of the test suite in detail . |
| Outcome: | The proposed test suite is based on linguistically motivated categories and phenomena and semi-automatic evaluation is carried out with regular expressions. |
Thermostat: A Large Collection of NLP Model Explanations and Analysis Tools (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Arras et al. (2016): explainability methods are perceived as opaque due to their complexity. |
| Approach: | They propose to use model explanations and analysis tools to facilitate research . they use a dataset that took 10k GPU hours to compile and analyse . |
| Outcome: | Thermostat allows easy access to over 200k explanations for state-of-the-art models . dataset took over 10k GPU hours (> one year) to compile; saves time . |
Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient’s Perspective (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that the class labels of german documents containing ADRs are imbalanced . clinical trials and physicians prescribing medications cannot cover every potential use case. |
| Approach: | They propose to use binary annotated documents from a german patient forum to detect ADRs. |
| Outcome: | The proposed model achieves an F1 score of 37.52 for the positive class on the German patient forum. |
Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas (2026.findings-acl)
Copied to clipboard
Lukas Stähelin, Veronika Solopova, Max Upravitelev, David Kaplan, Premtim Sahitaj, Ariana Sahitaj, Charlott Jakob, Sebastian Möller, Vera Schmitt
| Challenge: | Propaganda detection in social media is challenging due to noisy, short texts and low annotation agreements. |
| Approach: | They propose a new intent-focused taxonomy of propaganda techniques and compare it against an established, higher-agreement schema. |
| Outcome: | The proposed taxonomy outperforms existing models and reveals methodological differences hidden in base models. |
Using Neural Machine Translation Methods for Sign Language Translation (2022.acl-srw)
Copied to clipboard
| Challenge: | Sign languages are the main medium of exchanging information for the deaf and hard of hearing. |
| Approach: | They propose to use two NMT architectures to train models on parallel German Sign Language corpora . they achieve substantial improvement in BLEU scores for the models trained on the two corporales . |
| Outcome: | The proposed models achieve significant improvements on the two corpora trained on the german sign language . the proposed models outperform the models trained on both corporales . |
Towards a Reliable and Robust Methodology for Crowd-Based Subjective Quality Assessment of Query-Based Extractive Text Summarization (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing number of documents are needed for multi-document summarization. |
| Approach: | They propose crowdsourcing to evaluate intrinsic and extrinsic quality of extractive text summaries . they conduct intensive comparative crowdsourcing and laboratory experiments . |
| Outcome: | The proposed crowdsourcing task evaluates intrinsic and extrinsic quality of extractive text summaries. |
MultiTACRED: A Multilingual Version of the TAC Relation Extraction Dataset (2023.acl-long)
Copied to clipboard
| Challenge: | Relation extraction (RE) is a fundamental task in information extraction, but its extension to multilingual settings is hindered by the lack of supervised resources comparable in size to large English datasets. |
| Approach: | They propose a dataset to analyze relation extraction (RE) in multilingual settings . they find machine translation is a viable strategy to transfer RE instances . |
| Outcome: | The proposed dataset covers 12 typologically diverse languages from 9 language families and is compared with existing datasets. |
MuLVE, A Multi-Language Vocabulary Evaluation Data Set (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing systems for vocabulary evaluation are based on simple rules and do not account for real-life user learning data. |
| Approach: | They propose to use real-life user vocabulary learning data to evaluate vocabulary . they use language learning data from a phase6 vocabulary trainer to generate a multilingual data set for vocabulary evaluation. |
| Outcome: | The proposed data set provides outstanding results with 95.5 accuracy and F2-score. |
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models. |
| Approach: | They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts. |
| Outcome: | The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers. |
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages (2024.lrec-main)
Copied to clipboard
Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas, Tomohiro Nishiyama, Sebastian Möller, Eiji Aramaki, Yuji Matsumoto, Roland Roller, Pierre Zweigenbaum
| Challenge: | Existing clinical corpora mostly revolves around scientific articles in English . existing literature is limited to only a few scientific articles . |
| Approach: | They propose to use user-generated data sources to uncover adverse drug reactions . existing clinical corpora mostly revolves around scientific articles in english . authors provide statistics to highlight certain challenges associated with the corpus . |
| Outcome: | The proposed corpus includes 12 entity types, four attribute types, and 13 relation types . it provides strong baselines for extracting entities and relations between entities . |
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems (2025.findings-emnlp)
Copied to clipboard
Qianli Wang, Tatiana Anikina, Nils Feldhus, Simon Ostermann, Fedor Splitt, Jiaao Li, Yoana Tsoneva, Sebastian Möller, Vera Schmitt
| Challenge: | Current ConvXAI systems are based on intent recognition to accurately identify the user’s desired intention and map it to an explainability method. |
| Approach: | They propose a multilingual extension of the CoXQL dataset spanning five typologically diverse languages, including one low-resource language. |
| Outcome: | The proposed model enables multilingual generalization in a multilingual dataset spanning five typologically diverse languages, including one low-resource language. |
InterroLang: Exploring NLP Models and Datasets through Dialogue-based Explanations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on NLP explainability methods lacks a dialogue-based interpretability framework that can convey faithful explanations in human-understandable terms. |
| Approach: | They adapt the conversational explanation framework TalkToModel to the NLP domain and add new NLP-specific operations such as free-text rationalization to illustrate its generalizability. |
| Outcome: | The proposed framework can be used to explain models on three NLP tasks and is generalizable to different datasets, use cases and models. |