Papers by Nikolas Vitsakis
The Only Way is Ethics: A Guide to Ethical Research with Large Language Models (2025.coling-main)
Copied to clipboard
Eddie L. Ungless, Nikolas Vitsakis, Zeerak Talat, James Garforth, Bjorn Ross, Arno Onken, Atoosa Kasirzadeh, Alexandra Birch
| Challenge: | Existing literature on the ethical aspects of large language models (LLMs) is lacking a single practical guide on the subject. |
| Approach: | They propose to translate ethics literature into concrete recommendations for computer scientists by presenting an open and living resource for NLP practitioners and those tasked with evaluating the ethical implications of others’ work. |
| Outcome: | The proposed guide is an open and living resource for NLP practitioners and those tasked with evaluating the ethical implications of others’ work. |
Voices in a Crowd: Searching for clusters of unique perspectives (2024.emnlp-main)
Copied to clipboard
| Challenge: | Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata. |
| Approach: | They propose a framework that trains models without encoding annotator metadata and creates clusters of similar opinions, that are called voices. |
| Outcome: | The proposed framework captures minority perspectives based on demographic factors in two distinct datasets while also capturing majority perspectives. |
Re-examining Sexism and Misogyny Classification with Annotator Attitudes (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets for content moderation fail to capture plurality of possible annotator perspectives or ensure representation of affected groups. |
| Approach: | They examine the relationship between annotator identities and attitudes and the responses they give to two GBV labelling tasks. |
| Outcome: | The results show that higher Right Wing Authoritarianism scores are associated with a higher propensity to label text as sexist . higher scores are also associated with negative attitudes towards sexism and neosexist attitudes . |
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent Vision and Language models have shown impressive performance across benchmarks . however, frontier models lack cultural awareness and can affect global cultural diversity . |
| Approach: | They propose a visual question answering benchmark to probe the knowledge of culture-specific concepts and evaluate the capacity for cultural adaptation through contextual information. |
| Outcome: | The proposed model shows large performance disparities between culture-specific and common concepts in the parametric setting. |
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks (2024.emnlp-main)
Copied to clipboard
| Challenge: | Evaluating generalisation capabilities of multimodal models based solely on performance on out-of-distribution data fails to capture their true robustness . proposed framework examines the role of instructions and inputs in generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
| Approach: | They propose a framework that examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
| Outcome: | The proposed framework examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
AInterviewer: A Platform for Designing and Conducting AI-led Qualitative Interviews (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing systems rely on proprietary LLMs, which compromise reproducibility and data security. |
| Approach: | They propose a platform that combines controlled question administration of survey software with the flexibility of LLMs. |
| Outcome: | AInterviewer combines controlled question administration of survey software with flexibility of LLMs. |