Papers by Bertie Vidgen

13 papers
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)

Copied to clipboard

Challenge: Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect .
Approach: They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks .
Outcome: The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis .
HateCheck: Functional Tests for Hate Speech Detection Models (2021.acl-long)

Copied to clipboard

Challenge: Hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score.
Approach: They propose a suite of functional tests for hate speech detection models that measure model performance on held-out test data and then craft test cases to validate their quality.
Outcome: The proposed tests show that the proposed models perform poorly on a small set of widely-used hate speech datasets.
Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate (2022.naacl-main)

Copied to clipboard

Challenge: Existing models for detecting hate expressed with emojis have weaknesses when used for sensitive applications such as content moderation.
Approach: They propose a test suite of 3,930 short-form statements that evaluates hateful language expressed with emoji.
Outcome: The proposed model performs better on emoji-based hate while maintaining strong performance on text-only hate.
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)

Copied to clipboard

Challenge: Dynabench is an open-source platform for dynamic dataset creation and model benchmarking.
Approach: They propose an open-source platform for dynamic dataset creation and model benchmarking.
Outcome: The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios.
An Expert Annotated Dataset for the Detection of Online Misogyny (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies have found that misogynistic content is pervasive on some Reddit communities, but a training dataset for misogorical classification has not been created with the data.
Approach: They propose a hierarchical taxonomy and an expert labelled dataset to enable automatic classification of online misogynistic content.
Outcome: The proposed taxonomy and an expert labelled dataset are made freely available for future research.
Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks (2022.naacl-main)

Copied to clipboard

Challenge: Labelled data is the foundation of most natural language processing tasks, but there are valid beliefs about what the correct data labels should be.
Approach: They propose two contrasting paradigms for data annotation that encourage annotator subjectivity . they propose a descriptive paradigm that allows for the surveying and modelling of different beliefs .
Outcome: The proposed paradigms encourage annotator subjectivity, while the prescriptive paradigm discourages it.
Handling and Presenting Harmful Text in NLP Research (2022.findings-emnlp)

Copied to clipboard

Challenge: Text data can pose a risk of harm, but the risks remain unresolved in the NLP community.
Approach: They propose an analytical framework categorising harms on three axes: harm type, whether harm sought as a feature of research design, whether harmful content is encountered when working on unrelated problems, and who it affects .
Outcome: The proposed framework categorises harms on three axes: harm type, whether harm sought as feature of research design, and whether harmful content is encountered when working on unrelated problems.
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values (2023.emnlp-main)

Copied to clipboard

Challenge: Incorporating human feedback into Large Language Models is a welcome development, but it introduces new biases and challenges.
Approach: They propose to survey 95 articles that use human feedback to steer, guide or tailor the behaviours of large language models.
Outcome: The proposed approaches are based on 95 articles primarily from the ACL and arXiv repositories and highlight five unresolved conceptual and practical challenges.
Deciphering Implicit Hate: Evaluating Automated Detection Algorithms for Multimodal Hate (2021.findings-acl)

Copied to clipboard

Challenge: Imlicit hate content has unusual syntax, polysemic words, and fewer markers of prejudice, e.g., slurs . multimodal content is harder to detect than unimodal content, such as memes .
Approach: They evaluate the role of semantic and multimodal context for detecting implicit and explicit hate . they find that all models perform better on content with full annotator agreement .
Outcome: The proposed model outperforms other models on implicit and explicit hate detection tasks because of its lower propensity towards false positives.
LMUNIT: Fine-grained Evaluation with Natural Language Unit Tests (2025.findings-emnlp)

Copied to clipboard

Challenge: Using natural language unit tests, language models are costly and noisy, and automated metrics provide only coarse, difficult-to-interpret signals.
Approach: They propose a paradigm that decomposes response quality into explicit, testable criteria and a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales.
Outcome: The proposed paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows.
Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection (2021.acl-long)

Copied to clipboard

Challenge: Detecting online hate speech has proven difficult and concerns raised about performance, robustness, generalisability and fairness of stateof-the-art models.
Approach: They propose a human-and-model-in-the-loop process for dynamically generating datasets and training better performing hate detection models.
Outcome: The proposed model improves on a dataset of 40,000 hateful entries . the model is harder for annotators to trick and better on HateCheck .
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)

Copied to clipboard

Challenge: Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools.
Approach: They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations.
Outcome: The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are now being used by millions of people across the world.
Approach: They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way.
Outcome: The proposed test suite identifies eXaggerated Safety behaviours in a systematic way.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations