Papers by Simon Ostermann

22 papers
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge (L18-1)

Copied to clipboard

Challenge: Various approaches for script knowledge extraction and processing have been proposed in recent years.
Approach: They propose a dataset to evaluate natural language understanding approaches based on commonsense knowledge.
Outcome: The proposed dataset provides test cases for the broader natural language understanding community.
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Understanding and controlling behavior of large language models (LLMs) is an important topic in multilingual NLP.
Approach: They propose a lightweight parallel-question benchmark for evaluating language-forcing behavior in large language models across 32 languages.
Outcome: The proposed benchmark measures language steering in 32 languages across 32 languages.
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages (2025.acl-srw)

Copied to clipboard

Challenge: Low-resource languages (LRLs) face significant challenges in natural language processing due to limited data.
Approach: They evaluate adapter-based methods for adapting mLMs to low-resource languages . they use unstructured text and structured knowledge from ConceptNet to evaluate adapters .
Outcome: The proposed methods outperform large language models and LLaMA-3 and deepSeek-R1 models on low training data.
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multilingual anaphora resolution include images and video inputs.
Approach: They propose to include multimodal information in the form of images in anaphora resolution tasks.
Outcome: The proposed approach improves resolution by 10% for unseen languages.
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem (2025.coling-main)

Copied to clipboard

Challenge: Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions.
Approach: They propose a role-modeling approach that employs two LLMs as generator and critic to generate and refine NLEs.
Outcome: The proposed model outperforms self-refine and can perform with less powerful LLMs.
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing systems based on large language models (LLMs) are more precise and reliable in identifying users’ intentions, but the recognition of intents still presents a challenge in the case of ConvXAI, since little training data exist and the domain is highly specific.
Approach: They propose to use a dataset in the NLP domain for user intent recognition in ConvXAI to improve parsing performance.
Outcome: The proposed system outperforms existing methods and improves on existing ones.
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models.
Approach: They evaluate the performance of large language models and their generation strategies in 11 different languages using 3 NLP tasks and 4 open-source LLMs.
Outcome: The proposed generation strategies and their combinations yield strong results across 11 languages, including several extremely low-resource ones.
Common European Language Data Space (2024.lrec-main)

Copied to clipboard

Challenge: the Common European Language Data Space (LDS) is an integral part of the EU data strategy, which aims at developing a single market for data.
Approach: the Common European Language Data Space (LDS) is an integral part of the EU data strategy . its decentralised technical infrastructure and governance scheme are currently being developed by the LDS project .
Outcome: the Common European Language Data Space (LDS) is an integral part of the EU data strategy, which aims at developing a single market for data.
GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge (2025.findings-naacl)

Copied to clipboard

Challenge: Contextualized word embeddings are available for many languages, but their coverage is limited for low resourced languages.
Approach: They propose a method that integrates multilingual graph knowledge into the embeddings to make them green.
Outcome: The proposed method outperforms state-of-the-art embeddings on lexical similarity task while being parameter-free at inference time.
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting.
Approach: They propose a feature-level mechanistic analysis methodology to probe RL generalization using a controlled experimental setup.
Outcome: The proposed method identifies a compact, task-agnostic set of features that directly mediate generalization across diverse tasks.
Find-2-Find: Multitask Learning for Anaphora Resolution and Object Localization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing systems require large number of accurate annotations, such as image-level labels and location-level labeling.
Approach: They propose a joint anaphora resolution and object localization dataset targeting visual-linguistic ambiguity.
Outcome: The proposed framework improves visual-linguistic alignment and object localization with one joint model compared to a strong single-task baseline.
Commonsense Inference in Natural Language Processing (COIN) - Shared Task Report (D19-60)

Copied to clipboard

Challenge: The workshop on Commonsense Inference in NLP (COIN) evaluated text understanding systems' ability to draw inferences about facts that are not mentioned in the text, but that are assumed to be common ground.
Approach: They propose to use commonsense knowledge to evaluate systems' ability to answer questions/queries about a text.
Outcome: The proposed tasks evaluated systems in two contexts: Commonsense Inference and Commonsensible Inference.
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for counterfactual examples are lacking for many tasks.
Approach: They propose a faithful approach for leveraging important words from feature attribution methods to generate counterfactual examples in a zero-shot setting.
Outcome: The proposed framework outperforms state-of-the-art frameworks on many tasks.
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Captioning (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics fail to evaluate factual correctness in procedural video captions . Existing metrics rely on lexical overlap or holistic semantic similarity, but miss role-specific omissions resulting in hallucinations .
Approach: They propose a role-aware, fact-level evaluation framework that distinguishes conceptual facts from contextual facts.
Outcome: Experiments show that state-of-the-art captioning models produce fluent but incomplete descriptions with systematic errors.
Assessing Web Search Credibility and Response Groundedness in Chat Assistants (2026.eacl-long)

Copied to clipboard

Challenge: Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat.
Approach: They propose a method for evaluating assistants’ web search behavior focusing on source credibility and the groundedness of responses with respect to cited sources.
Outcome: The proposed method focuses on source credibility and the groundedness of responses with respect to cited sources.
Large Language Models for Multilingual Previously Fact-Checked Claim Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study evaluates large language models for multilingual previously fact-checked claim detection . authors assess seven LLMs across 20 languages in monolingual and cross-lingual settings .
Approach: They evaluate large language models for multilingual previously fact-checked claim detection . they find they perform well for high-resource languages, struggle with low-resourced languages .
Outcome: The proposed model performs well for high-resource languages, but struggle with low-resourced languages.
HybridBERT - Making BERT Pretraining More Efficient Through Hybrid Mixture of Attention Mechanisms (2024.naacl-srw)

Copied to clipboard

Challenge: Pretrained transformer-based language models have produced state-of-the-art performance in most natural language understanding tasks.
Approach: They propose two hybrid architectures that combine self-attention and additive attention mechanisms with sub-layer normalization to achieve double the pretraining accuracy of a vanilla-BERT baseline.
Outcome: The proposed architectures outperform BERT-base on two downstream tasks while accelerating inference.
Mapping Texts to Scripts: An Entailment Study (L18-1)

Copied to clipboard

Challenge: Script knowledge is crucial for text understanding systems, providing a basis for commonsense inference.
Approach: They propose to map event mentions in a text to script events using crowdsourced event descriptions.
Outcome: The proposed model improves the performance of text-to-script mapping systems by integrating paraphrase sets with crowdsourced event descriptions.
From Weights to Activations: Is Steering the Next Frontier of Adaptation? (2026.acl-long)

Copied to clipboard

Challenge: Pre-trained large language models are the basis of a wide range of NLP tasks.
Approach: They propose to use parameter updates and parameter-efficient adaptation to modify behavior of large language models.
Outcome: The proposed method enables local and reversible behavioral change without parameter updates.
Soft Language Prompts for Language Transfer (2025.naacl-long)

Copied to clipboard

Challenge: Cross-lingual knowledge transfer, especially between high- and low-resource languages, remains challenging in natural language processing.
Approach: They propose to combine language-specific adapters and soft prompts to enhance cross-lingual transfer by parameter-efficient fine-tuning methods.
Outcome: The proposed methods outperform language adapters and soft prompts in 16 languages and 10 low-resource languages.
Only for the Unseen Languages, Say the Llamas: On the Efficacy of Language Adapters for Cross-lingual Transfer in English-centric LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Most state-of-the-art large language models (LLMs) are trained mainly on English data, limiting their effectiveness on non-English, especially low-resource, languages.
Approach: They train language adapters for 13 languages and evaluate their effectiveness on downstream tasks using either task adapters or in-context learning.
Outcome: The proposed language adapters improve performance for languages not seen during pretraining, but provide negligible benefit for seen languages.
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Current ConvXAI systems are based on intent recognition to accurately identify the user’s desired intention and map it to an explainability method.
Approach: They propose a multilingual extension of the CoXQL dataset spanning five typologically diverse languages, including one low-resource language.
Outcome: The proposed model enables multilingual generalization in a multilingual dataset spanning five typologically diverse languages, including one low-resource language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations