Papers by Roman Klinger

34 papers
iPrOp: Interactive Prompt Optimization for Large Language Models with a Human in the Loop (2025.acl-srw)

Copied to clipboard

Challenge: Prompt engineering has made significant contributions to the era of large language models, yet its effectiveness depends on the skills of a prompt author.
Approach: They propose a novel approach to prompt optimization that bridges manual prompt engineering and automatic prompt optimization by providing task-specific guidance.
Outcome: The proposed approach bridges manual prompt engineering and automatic prompt optimization while offering users the flexibility to assess evolving prompts.
Frowning Frodo, Wincing Leia, and a Seriously Great Friendship: Learning to Classify Emotional Relationships of Fictional Characters (N19-1)

Copied to clipboard

Challenge: Existing literature analysis does not focus on roles of characters or on relationships between them.
Approach: They propose to combine emotion and character identification into a unified framework for character network extraction from fictional texts.
Outcome: The proposed task is based on fan-fiction short stories and is able to predict emotion relations in the extracted network graph.
Dealing with Controversy: An Emotion and Coping Strategy Corpus Based on Role Playing (2024.findings-emnlp)

Copied to clipboard

Challenge: Psychological studies aim at explaining internal mechanisms of emotions, while computational studies simplify them into labels.
Approach: They propose to treat emotions as strategies to cope with salient situations . they introduce a task of coping identification and a corpus constructed via role-playing .
Outcome: The proposed method allows to investigate the link between emotions and behavior, which also emerges in language.
Which Demographics do LLMs Default to During Annotation? (2025.acl-long)

Copied to clipboard

Challenge: Demographics and cultural background of annotators influence the labels they assign in text annotation.
Approach: They examine the attributes of human annotators LLMs inherently mimic and compare them to demographic-conditioned prompts and placebo-conditioned ones.
Outcome: The proposed model incorporates demographics and cultural background into the output of the large language models (LLMs) to evaluate which attributes of human annotators LLMs inherently mimic.
CoVERT: A Corpus of Fact-checked Biomedical COVID-19 Tweets (2022.lrec-1)

Copied to clipboard

Challenge: Existing fact-checking resources cover COVID-19 related information in news, but there is no dataset providing fact- checked COVId-19 related tweets with detailed annotations for biomedical entities, relations and relevant evidence.
Approach: They propose a fact-checked corpus of tweets with annotations for biomedical entities, relations and relevant evidence for COVID-19 related tweets.
Outcome: The proposed dataset provides fact-checked COVID-19 related tweets with detailed annotations for biomedical entities, relations and relevant evidence.
What Makes Medical Claims (Un)Verifiable? Analyzing Entity and Relation Properties for Fact Verification (2024.eacl-long)

Copied to clipboard

Challenge: Existing studies show that identifying verifiable claims is difficult, whereas identifying unverifiably claims is more challenging.
Approach: They hypothesize that breaking down claims into smaller units increases our understanding which properties impact verifiability.
Outcome: The proposed corpus of evidence is based on the first corpus for scientific fact verification annotated with subject–relation–object triplets, evidence documents, and fact-checking verdicts.
Recovering Patient Journeys: A Corpus of Biomedical Entities and Relations on Twitter (BEAR) (2022.lrec-1)

Copied to clipboard

Challenge: Existing medical social media corpora focus on a small set of entities and relations . existing text mining and information extraction methods focus on scientific text generated by researchers but their access to individual patient experiences or patient-doctor interactions is limited.
Approach: The dataset consists of 2,100 medical tweets with approx. 6,000 entities and 2,200 relations.
Outcome: The proposed dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations.
Understanding Fine-grained Distortions in Reports of Scientific Findings (2024.findings-acl)

Copied to clipboard

Challenge: a fine-grained understanding of how scientific findings are reported is crucial, says a new study . a recent study found that tweets distort scientific findings more often than news reports .
Approach: They propose to annotate 1,600 scientific findings from academic papers paired with corresponding tweets . they also establish baselines for automatically detecting these characteristics .
Outcome: The proposed method outperforms few-shot prompting in detecting distortions in unpaired data.
How Entangled is Factuality and Deception in German? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing research on deception detection and fact checking conflates factual accuracy with truthfulness . a belief-based deception framework defines texts as deceptive when there is a mismatch between what people say and what they truly believe .
Approach: They assess if presumed patterns of deception generalize to German language texts . they gauge the impact of deceptiveness on the downstream task of fact checking .
Outcome: The proposed framework disentangles deception when there is a mismatch between what people say and what they truly believe . the proposed framework does not find any correlation with established cues of deception .
Appraisal Theories for Emotion Classification in Text (2020.coling-main)

Copied to clipboard

Challenge: Automatic emotion categorization is based on textual units assigned to an emotion from a predefined inventory, for instance following the basic emotion classes proposed by Paul Ekman (1999) or Plutchik (2001).
Approach: They propose to make automatic emotion categorization explicit by following theories of cognitive appraisal of events and show their potential for emotion classification when being encoded in classification models.
Outcome: The proposed models improve the classification of discrete emotion categories by using appraisal dimension assignments in event descriptions.
MOPO: Multi-Objective Prompt Optimization for Affective Text Generation (2025.coling-main)

Copied to clipboard

Challenge: Using multi-objective prompt optimization, users can choose the most appropriate prompt for their context.
Approach: They propose a multi-objective prompt optimization methodology that optimizes prompts according to multiple objectives.
Outcome: The proposed method improves performance by 15 pp across all objectives with a minimal loss (1–2 pp)
Projecting Embeddings for Domain Adaption: Joint Modeling of Sentiment Analysis in Diverse Domains (C18-1)

Copied to clipboard

Challenge: Existing domain adaptation methods for sentiment analysis are sensitive to domain differences, resulting in classifiers that perform poorly on new domains.
Approach: They propose a domain adaptation problem as an embedding projection task using two mono-domain embeddable spaces and a bi-domain space to project across domains and predict sentiment.
Outcome: The proposed model performs better on domains similar to state-of-the-art methods while requiring longer training times.
DERE: A Task and Domain-Independent Slot Filling Framework for Declarative Relation Extraction (D18-2)

Copied to clipboard

Challenge: Comparability of models across tasks is lacking in most machine learning systems for natural language processing.
Approach: They propose a framework for declarative specification and compilation of template-based information extraction that uses a generic specification language for the task and for data annotations in terms of spans and frames.
Outcome: The proposed framework enables representation of a large variety of natural language processing tasks.
Who Feels What and Why? Annotation of a Literature Corpus with Semantic Roles of Emotions (C18-1)

Copied to clipboard

Challenge: Emotion analysis and classification is a challenging task which has been tackled with relatively straight-forward approaches.
Approach: They propose to annotate emotion trigger phrases and entities in the roles of experiencers, targets, and causes of the emotion in literature by Project Gutenberg.
Outcome: The proposed corpus supports qualitative literary studies and digital humanities.
Lost in Back-Translation: Emotion Preservation in Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: MT is used to support human-to-human communication across languages, but it is unclear whether it can translate the non-propositional level of emotions.
Approach: They propose to use a re-ranking approach to change emotions to reverse this tendency . they find that emotions are toned down or amplified through linguistic changes .
Outcome: The proposed model can be used to change emotions, and it can be applied to other languages.
EmoProgress: Cumulated Emotion Progression Analysis in Dreams and Customer Service Dialogues (2024.lrec-main)

Copied to clipboard

Challenge: Emotion analysis often involves categorization of isolated textual units, but these are parts of longer discourses, like dialogues or stories.
Approach: They propose a novel annotation setup for emotion categorization corpora that allows to annotate the emotion up to the annotated sentence.
Outcome: The proposed annotation setup allows to answer the question which emotion is presumably experienced at a specific moment in time.
Token Sequence Labeling vs. Clause Classification for English Emotion Stimulus Detection (2020.starsem-1)

Copied to clipboard

Challenge: Emotion stimulus detection is the task of finding the cause of an emotion in a textual description.
Approach: They propose to evaluate whether clause classification or token sequence labeling is better for emotion stimulus detection in English.
Outcome: The proposed framework compares clause classification and token sequence labeling on four English datasets.
Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection (2026.eacl-long)

Copied to clipboard

Challenge: Existing computational approaches focus on logical structures of fallacies and argumentation schemes, ignoring the emotional dimension of argumentation.
Approach: They propose to use large language models to systematically change emotional appeals in fallacious arguments by using a computational approach.
Outcome: The proposed method reduces fallacy detection by 14.5% on average on human arguments with enjoyment over fear or sadness.
GoodNewsEveryone: A Corpus of News Headlines Annotated with Emotions, Semantic Roles, and Reader Perception (2020.lrec-1)

Copied to clipboard

Challenge: Fewer studies address emotions as a phenomenon to be tackled with structured learning, which can be explained by the lack of relevant datasets.
Approach: They propose to annotate 5000 English news headlines with their associated emotions, the corresponding emotion experiencers and textual cues, related emotion causes and targets, and the reader’s perception of the emotion of the headline.
Outcome: The proposed method enables further research on emotion classification, emotion intensity prediction, emotion cause detection and supports qualitative studies.
An Analysis of Annotated Corpora for Emotion Classification in Text (C18-1)

Copied to clipboard

Challenge: Several datasets have been annotated and published for classification of emotions.
Approach: They aggregated emotion corpora in a common file format with a shared annotation schema . they perform cross-corpus classification experiments to gain insight and a better understanding of differences .
Outcome: The proposed model can be trained on a subset of corpora, but not on all corporata.
PO-EMO: Conceptualization, Annotation, and Modeling of Aesthetic Emotions in German and English Poetry (2020.lrec-1)

Copied to clipboard

Challenge: a new study shows that literature enables engagement in a broader range of complex and subtle emotions.
Approach: They propose to use multiple emotion labels to capture mixed emotions in poetry . they evaluate an annotation experiment with experts and crowdsourcing .
Outcome: The proposed method shows that identifying aesthetic emotions is challenging in the German subset.
Can Factual Statements Be Deceptive? The DeFaBel Corpus of Belief-based Deception (2024.lrec-main)

Copied to clipboard

Challenge: if a person firmly believes in a non-factual statement, there is no inherent intention to deceive.
Approach: They propose to use the DeFaBel corpus to study the relationship between deception and factuality based on belief to generate arguments supporting statements .
Outcome: The DeFaBel corpus contains 1031 texts in german, out of which 643 are deceptive and 388 are non-deceptive.
Natural Language Inference Prompts for Zero-shot Emotion Classification in Text across Corpora (2022.coling-1)

Copied to clipboard

Challenge: Existing models for textual emotion classification depend on domain and application scenario and need to be predefined . a natural language inference model with a flexible set of labels is difficult to develop .
Approach: They propose to use the paradigm of zero-shot learning as a natural language inference task to generate a model with a flexible set of labels.
Outcome: The proposed model is more robust across corpora than individual prompts and shows similar performance to the best prompt for a particular corpus.
Bilingual Sentiment Embeddings: Joint Projection of Sentiment Across Languages (P18-1)

Copied to clipboard

Challenge: Existing approaches to sentiment analysis in low-resource languages lack annotated corpora or do not capture sentiment information.
Approach: They propose a model that represents sentiment in a source and target language without annotated corpus.
Outcome: The proposed model outperforms state-of-the-art methods on four out of six setups and captures complementary information to machine translation.
Automatic Section Recognition in Obituaries (2020.lrec-1)

Copied to clipboard

Challenge: Obituaries contain information about people’s values across times and cultures, which makes them useful for exploring cultural history.
Approach: They propose to use a convolutional neural network to recognize these sections in obituaries to improve their annotation.
Outcome: The proposed model outperforms bag-of-words and embedding-based BiLSTMs and BiLStm-CRFs with a micro F1 = 0.81.
x-enVENT: A Corpus of Event Descriptions with Experiencer-specific Emotion and Appraisal Annotations (2022.lrec-1)

Copied to clipboard

Challenge: Emotion classification is often formulated as the task to categorize texts into a predefined set of emotion classes.
Approach: They propose that a classification setup for emotion analysis should be performed in an integrated manner, including the different semantic roles that participate in an emotion episode.
Outcome: The proposed method reveals patterns in the co-occurrence of people’s emotions in interaction.
SANTO: A Web-based Annotation Tool for Ontology-driven Slot Filling (P18-4)

Copied to clipboard

Challenge: SANTO is an annotation tool designed for complex relation extraction tasks . a subset of information extraction tasks can be typed n-ary relation extraction or slot filling .
Approach: They propose a domain-adaptive annotation tool for complex slot filling tasks . SANTO enables fast and clearly structured annotation for multiple users in parallel .
Outcome: The proposed tool can be used for slot filling tasks and import and export procedures of standard formats enable interoperability with external sources and tools.
Embarrassingly Simple Performance Prediction for Abductive Natural Language Inference (2022.naacl-main)

Copied to clipboard

Challenge: a method for learning an NLI model is time-consuming and resource-intensive, but it can save time and resources.
Approach: They propose a method for predicting model performance without fine-tuning it . they compare sentence embeddings with cosine similarity to classifiers .
Outcome: The proposed method can save time and resources by comparing pre-trained models to real-world datasets.
“You are an expert annotator”: Automatic Best–Worst-Scaling Annotations for Emotion Intensity Modeling (2024.naacl-long)

Copied to clipboard

Challenge: Large language models mitigate the issue with automatic corpus labeling methods, but there is no work on automating annotations for continuous labels.
Approach: They propose to use a transformer regressor to automate emotion intensity predictions and compare rating scale predictions with best–worst scaling.
Outcome: The proposed method performs better on rating scale annotation tasks than on comparative annotation tasks.
Emotion Analysis from Texts (2023.eacl-tutorials)

Copied to clipboard

Challenge: Emotion analysis in text is a field of research that encompasses a set of various natural language processing tasks.
Approach: This tutorial provides an overview of research from emotion psychology . it discusses the use cases of emotion analysis in text, their societal impact and ethical considerations .
Outcome: This paper provides an overview of research from emotion psychology which sets the ground for choosing adequate NLP methodology.
Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts (2025.acl-long)

Copied to clipboard

Challenge: Accurate modeling of subjective phenomena requires data annotated with authors’ intentions.
Approach: They collect study-created and genuine social media posts labeled for emotion and compare them on several dimensions, including model performance.
Outcome: The results show that study-created posts are longer, rely more on text and less on images for emotion expression, and focus more on emotion-prototypical events.
Crowdsourcing and Validating Event-focused Emotion Corpora for German and English (P19-1)

Copied to clipboard

Challenge: Existing studies on automatic recognition of emotions in text have achieved promising results, but there is a shortage of resources for non-English languages, with few exceptions, like Chinese.
Approach: They propose to use a crowdsourced German emotion corpus to build a corpus similar to the English ISEAR emotion dataset.
Outcome: The proposed model performs well in German and English, but lacks the resources for non-English languages.
Adversarial Training for Satire Detection: Controlling for Confounding Variables (N19-1)

Copied to clipboard

Challenge: Existing methods for satire detection focus on satirical news based on article sources . satiric news are written with the aim of mimicking regular news in diction .
Approach: They propose a model for satire detection with an adversarial component to control for the confounding variable of publication source.
Outcome: The proposed model improves generalization performance to unseen publications with an adversarial component.
Dissecting Span Identification Tasks with Performance Prediction (2020.emnlp-main)

Copied to clipboard

Challenge: Span identification tasks are a staple of applied NLP, but there is little insight on how their properties influence their difficulty.
Approach: They propose to build a model to predict span ID performance for unseen span ID tasks that can support architecture choices.
Outcome: The proposed model predicts span ID tasks for unseen span ID task in English, and the meta model predictable span ID performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations