Papers by Verena Rieser
Adversarial Textual Robustness on Visual Dialog (2023.findings-acl)
Copied to clipboard
| Challenge: | a recent study evaluated the robustness of visual dialog models against textual attacks. |
| Approach: | They aim to understand how multimodal input components contribute to robustness . they also evaluate how to generate adversarial test examples which fool the model . |
| Outcome: | The proposed model is more robust when it encodes dialog history than when it does not. |
Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity (D18-1)
Copied to clipboard
| Challenge: | Existing encoder-decoder models for open domain dialogue generate generic, uninformative, and non-coherent responses. |
| Approach: | They propose to introduce a measure of coherence as the GloVe embedding similarity between dialogue context and generated response to improve output diversity. |
| Outcome: | The proposed model improves on the OpenSubtitles corpus in terms of BLEU score and diversity metrics. |
OTTers: One-turn Topic Transitions for Open-Domain Dialogue (2021.acl-long)
Copied to clipboard
| Challenge: | a mixed-initiative dialogue system is often purely responsive, make abrupt transitions, or fail to take initiative. |
| Approach: | They propose a task to generate a "bridging" utterance connecting a new topic to the previous conversation turn. |
| Outcome: | The proposed task generates a "bridging" utterance connecting a new topic to the previous topic. |
Mirages. On Anthropomorphism in Dialogue Systems (2023.emnlp-main)
Copied to clipboard
| Challenge: | Automated dialogue systems are anthropomorphised by developers and personified by users. |
| Approach: | They propose to examine linguistic factors that contribute to the anthropomorphism of dialogue systems and the harms that can arise thereof. |
| Outcome: | The proposed systems are anthropomorphised and personified by users . linguistic factors can also be used to reinforce gender stereotypes and conceptions of acceptable language. |
MiRANews: Dataset and Benchmarks for Multi-Resource-Assisted News Summarization (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Current news summarization systems often contain 'extrinsic hallucinations', i.e. facts that are not present in the source document, which are often derived via world knowledge. |
| Approach: | They propose to use multiple supplementary resource documents to assist the task by pairing a single document with a human authored summary as the summary. |
| Outcome: | The proposed model reduces 55% of hallucinations when compared to single-document summarization models trained on the main article only. |
Risk-graded Safety for Handling Medical Queries in Conversational AI (2022.aacl-short)
Copied to clipboard
| Challenge: | Conversational AI systems can engage in unsafe behaviour when handling medical queries that could lead to death. |
| Approach: | They label medical queries with crowdsourced and expert annotations to identify the seriousness of the prompts and recognise the risk types posed by the responses. |
| Outcome: | The results suggest that these tasks can be automated, but caution should be exercised, as errors can potentially be very serious. |
The Dangers of trusting Stochastic Parrots: Faithfulness and Trust in Open-domain Conversational Question Answering (2023.findings-acl)
Copied to clipboard
Sabrina Chiesurin, Dimitris Dimakopoulos, Marco Antonio Sobrevilla Cabezudo, Arash Eshghi, Ioannis Papaioannou, Verena Rieser, Ioannis Konstas
| Challenge: | Empirical studies of dialogue have shown that people use different kinds of context-dependent linguistic behavior to indicate grounding, including use of fragments, ellipsis and pronominal reference. |
| Approach: | They propose to use open-domain question answering systems as test-bed for task based dialog generation and compare open- and closed-book models to test their hypothesis. |
| Outcome: | The proposed model parrots user input while providing an unfaithful response. |
Multitask Multimodal Prompted Training for Interactive Embodied Task Completion (2023.emnlp-main)
Copied to clipboard
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia
| Challenge: | Embodied MultiModal Agent (EMMA) is a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text generation. |
| Approach: | They propose an Embodied MultiModal Agent (EMMA) that uses a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text. |
| Outcome: | The proposed model performs on par with similar models on several VL benchmarks and sets a new state-of-the-art success rate on the Dialog-guided Task Completion (DTC) benchmark. |
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics (2025.findings-emnlp)
Copied to clipboard
Shravan Nayak, Mehar Bhatia, Xiaofeng Zhang, Verena Rieser, Lisa Anne Hendricks, Sjoerd Van Steenkiste, Yash Goyal, Karolina Stanczak, Aishwarya Agrawal
| Challenge: | CulturalFrames is a benchmark designed for rigorous human evaluation of cultural representation in visual generations. |
| Approach: | They propose to quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied by the prompt’s cultural context) cultural expectations. |
| Outcome: | The proposed model is based on 983 prompts, 3637 images and 10k human annotations from 10 countries and 5 socio-cultural domains. |
RankME: Reliable Human Ratings for Natural Language Generation (N18-2)
Copied to clipboard
| Challenge: | Existing studies have shown that human evaluation for natural language generation often suffers from inconsistent user ratings. |
| Approach: | They propose a rank-based magnitude estimation method which combines continuous scales and relative assessments to improve the reliability of human ratings. |
| Outcome: | The proposed method significantly improves the reliability and consistency of human ratings compared to traditional evaluation methods. |
Value Profiles for Encoding Human Variation (2025.emnlp-main)
Copied to clipboard
Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah Goodman, Verena Rieser
| Challenge: | Using value profiles and a steerable decoder model to estimate ratings is crucial for personalization, pluralistic model alignment, and computational social science. |
| Approach: | They propose to represent individuals using value profiles and a steerable decoder model to estimate ratings conditioned on a value profile or other rater information. |
| Outcome: | The proposed model interpretably changes ratings according to semantic profile differences and is well-calibrated. |
Fact-based Content Weighting for Evaluating Abstractive Summarisation (2020.acl-main)
Copied to clipboard
| Challenge: | Abstractive summarisation is notoriously hard to evaluate since word-overlap-based metrics are insufficient. |
| Approach: | They propose a new evaluation metric which is based on fact-level content weighting, relating the facts of the document to the facts in the summary. |
| Outcome: | The proposed evaluation metric is highly correlated to human perception and compares favourably to the recent manual highlight-based metric of Hardy et al. |
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Several studies discuss the potential harms and benefits of large language models (LLMs) large neural models can replicate and even amplify negative, stereotypical, and derogatory associations in the data. |
| Approach: | They propose to use a first aid kit to assess the safety of conversational AI in various settings . they propose several future directions and discuss ethical considerations . |
| Outcome: | The proposed tools can provide estimates of the relative safety of systems in various settings, but they still have several shortcomings. |
What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you think (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that human evaluations in NLP are under-powered because of two common factors: they treat ordinal data as interval data and operate under high variance settings. |
| Approach: | They propose to use ordinal mixed effects models to detect small differences between models, especially in high variance settings common in NLP evaluations of generated texts. |
| Outcome: | The proposed models detect small differences in high variance settings, especially in high-variance evaluations of generated texts. |
SLURP: A Spoken Language Understanding Resource Package (2020.emnlp-main)
Copied to clipboard
| Challenge: | Publicly available datasets for Spoken Language Understanding (SLU) are limited. |
| Approach: | They propose a publicly available SLU resource package that includes a multi-domain dataset in English spanning 18 domains. |
| Outcome: | The proposed dataset is bigger and more diverse than existing datasets. |
STAR: SocioTechnical Approach to Red Teaming Language Models (2024.emnlp-main)
Copied to clipboard
Laura Weidinger, John Mellor, Bernat Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Bergman, Mikel Rodriguez, Verena Rieser, William Isaac
| Challenge: | STAR is a sociotechnical framework that improves on current best practices for red teaming safety of large language models. |
| Approach: | They propose a sociotechnical framework that improves on current best practices for red teaming safety of large language models. |
| Outcome: | The proposed framework improves on current best practices for red teaming safety of large language models. |
History for Visual Dialog: Do we really need it? (2020.acl-main)
Copied to clipboard
| Challenge: | Recent studies have shown that dialog-based interaction grounded in visual information is not as effective as previous VQA tasks because of its dialog history. |
| Approach: | They propose a visual dialogue subset which explicitly encodes dialog history and a NDCG benchmark of 63%. |
| Outcome: | The proposed subset (VisdialConv) of the VisdialVal set achieves state-of-the-art performance on 72 % of the data. |
ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Abuse Detection in Conversational AI (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on abusive language towards conversational AI systems are not conclusive as they are not performed with live systems nor with real users due to the lack of reliable abuse detection tools. |
| Approach: | They propose to use a convAI dataset to account for the complexity of the task and to bench-mark existing models against this data. |
| Outcome: | The proposed model shows that abuse distribution is different compared to other datasets, with sexual tinted aggression towards the virtual persona of the systems. |
AggGen: Ordering and Aggregating while Generating (2021.acl-long)
Copied to clipboard
| Challenge: | AggGen is a data-to-text model which re-introduces two explicit sentence planning stages into neural data- to-text systems: input ordering and input aggregation. |
| Approach: | AggGen re-introduces two explicit sentence planning stages into neural data-to-text systems: input ordering and input aggregation. |
| Outcome: | AggGen is a data-to-text model which re-introduces two explicit sentence planning stages into neural data- to-text systems: input ordering and input aggregation. |