Papers by Sebastian Möller

20 papers
Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)

Copied to clipboard

Challenge: Evaluating translation models is a trade-off between effort and detail.
Approach: They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts.
Outcome: The proposed method exposes systematic differences between human and machine translations to human experts.
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem (2025.coling-main)

Copied to clipboard

Challenge: Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions.
Approach: They propose a role-modeling approach that employs two LLMs as generator and critic to generate and refine NLEs.
Outcome: The proposed model outperforms self-refine and can perform with less powerful LLMs.
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing systems based on large language models (LLMs) are more precise and reliable in identifying users’ intentions, but the recognition of intents still presents a challenge in the case of ConvXAI, since little training data exist and the domain is highly specific.
Approach: They propose to use a dataset in the NLP domain for user intent recognition in ConvXAI to improve parsing performance.
Outcome: The proposed system outperforms existing methods and improves on existing ones.
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation (2026.acl-long)

Copied to clipboard

Challenge: Large language models excel at generating English counterfactuals but their effectiveness in generating multilingual counterfacts remains unclear.
Approach: They conduct automatic evaluations on both directly generated and derived counterfactuals in six languages and find that cross-lingual perturbations follow common strategic principles.
Outcome: The proposed models show that translation-based counterfactuals offer higher validity than their directly generated counterparts, but still fall short of matching the quality of the original English counterf actuals.
From Witch’s Shot to Music Making Bones - Resources for Medical Laymen to Technical Language and Vice Versa (2020.lrec-1)

Copied to clipboard

Challenge: Information we share online unveils directly or indirectly information about our lifestyle and health situation.
Approach: They propose a dataset which annotates medical laymen and technical expressions in a patient forum and a set of medical synonyms and definitions.
Outcome: The proposed dataset annotates medical laymen and technical expressions in a patient forum along with a set of medical synonyms and definitions.
An Empirical Comparison of Question Classification Methods for Question Answering Systems (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for Question Classification are monolingual, but they are not suitable for low-resourced languages.
Approach: They propose to classify the most recent methods in four different categories . they propose to use a low, medium, high, and very high level of dependency on external resources .
Outcome: The proposed method outperforms methods not suitable for low-resource languages.
Subjective Text Complexity Assessment for German (2022.lrec-1)

Copied to clipboard

Challenge: Often, readability is defined as how easily a written text is to read.
Approach: They propose to use a corpus of sentences provided by a German IT service provider to assess the readability of German text.
Outcome: The proposed model can predict complexity of German text by using linguistically motivated features.
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for counterfactual examples are lacking for many tasks.
Approach: They propose a faithful approach for leveraging important words from feature attribution methods to generate counterfactual examples in a zero-shot setting.
Outcome: The proposed framework outperforms state-of-the-art frameworks on many tasks.
A Linguistically Motivated Test Suite to Semi-Automatically Evaluate German–English Machine Translation Output (2022.lrec-1)

Copied to clipboard

Challenge: Using fine-grained evaluation techniques, translation outputs have become better and more fluent.
Approach: They propose a fine-grained test suite for the language pair German–English . they describe the creation and implementation of the test suite in detail .
Outcome: The proposed test suite is based on linguistically motivated categories and phenomena and semi-automatic evaluation is carried out with regular expressions.
Thermostat: A Large Collection of NLP Model Explanations and Analysis Tools (2021.emnlp-demo)

Copied to clipboard

Challenge: Arras et al. (2016): explainability methods are perceived as opaque due to their complexity.
Approach: They propose to use model explanations and analysis tools to facilitate research . they use a dataset that took 10k GPU hours to compile and analyse .
Outcome: Thermostat allows easy access to over 200k explanations for state-of-the-art models . dataset took over 10k GPU hours (> one year) to compile; saves time .
Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient’s Perspective (2022.lrec-1)

Copied to clipboard

Challenge: a recent study shows that the class labels of german documents containing ADRs are imbalanced . clinical trials and physicians prescribing medications cannot cover every potential use case.
Approach: They propose to use binary annotated documents from a german patient forum to detect ADRs.
Outcome: The proposed model achieves an F1 score of 37.52 for the positive class on the German patient forum.
Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas (2026.findings-acl)

Copied to clipboard

Challenge: Propaganda detection in social media is challenging due to noisy, short texts and low annotation agreements.
Approach: They propose a new intent-focused taxonomy of propaganda techniques and compare it against an established, higher-agreement schema.
Outcome: The proposed taxonomy outperforms existing models and reveals methodological differences hidden in base models.
Using Neural Machine Translation Methods for Sign Language Translation (2022.acl-srw)

Copied to clipboard

Challenge: Sign languages are the main medium of exchanging information for the deaf and hard of hearing.
Approach: They propose to use two NMT architectures to train models on parallel German Sign Language corpora . they achieve substantial improvement in BLEU scores for the models trained on the two corporales .
Outcome: The proposed models achieve significant improvements on the two corpora trained on the german sign language . the proposed models outperform the models trained on both corporales .
Towards a Reliable and Robust Methodology for Crowd-Based Subjective Quality Assessment of Query-Based Extractive Text Summarization (2020.lrec-1)

Copied to clipboard

Challenge: a growing number of documents are needed for multi-document summarization.
Approach: They propose crowdsourcing to evaluate intrinsic and extrinsic quality of extractive text summaries . they conduct intensive comparative crowdsourcing and laboratory experiments .
Outcome: The proposed crowdsourcing task evaluates intrinsic and extrinsic quality of extractive text summaries.
MultiTACRED: A Multilingual Version of the TAC Relation Extraction Dataset (2023.acl-long)

Copied to clipboard

Challenge: Relation extraction (RE) is a fundamental task in information extraction, but its extension to multilingual settings is hindered by the lack of supervised resources comparable in size to large English datasets.
Approach: They propose a dataset to analyze relation extraction (RE) in multilingual settings . they find machine translation is a viable strategy to transfer RE instances .
Outcome: The proposed dataset covers 12 typologically diverse languages from 9 language families and is compared with existing datasets.
MuLVE, A Multi-Language Vocabulary Evaluation Data Set (2022.lrec-1)

Copied to clipboard

Challenge: Existing systems for vocabulary evaluation are based on simple rules and do not account for real-life user learning data.
Approach: They propose to use real-life user vocabulary learning data to evaluate vocabulary . they use language learning data from a phase6 vocabulary trainer to generate a multilingual data set for vocabulary evaluation.
Outcome: The proposed data set provides outstanding results with 95.5 accuracy and F2-score.
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models.
Approach: They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts.
Outcome: The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers.
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing clinical corpora mostly revolves around scientific articles in English . existing literature is limited to only a few scientific articles .
Approach: They propose to use user-generated data sources to uncover adverse drug reactions . existing clinical corpora mostly revolves around scientific articles in english . authors provide statistics to highlight certain challenges associated with the corpus .
Outcome: The proposed corpus includes 12 entity types, four attribute types, and 13 relation types . it provides strong baselines for extracting entities and relations between entities .
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Current ConvXAI systems are based on intent recognition to accurately identify the user’s desired intention and map it to an explainability method.
Approach: They propose a multilingual extension of the CoXQL dataset spanning five typologically diverse languages, including one low-resource language.
Outcome: The proposed model enables multilingual generalization in a multilingual dataset spanning five typologically diverse languages, including one low-resource language.
InterroLang: Exploring NLP Models and Datasets through Dialogue-based Explanations (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work on NLP explainability methods lacks a dialogue-based interpretability framework that can convey faithful explanations in human-understandable terms.
Approach: They adapt the conversational explanation framework TalkToModel to the NLP domain and add new NLP-specific operations such as free-text rationalization to illustrate its generalizability.
Outcome: The proposed framework can be used to explain models on three NLP tasks and is generalizable to different datasets, use cases and models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations