Papers by Alessandra Urbinati
Are you sure? Measuring models bias in content moderation through uncertainty (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Language Model-based classifiers perpetuate racial and social biases in content moderation . et al., j. n. d., and j neil, e. c. (2005) measure the fairness of content moderated models . |
| Approach: | They propose an unsupervised approach that benchmarks models on their uncertainty . they use uncertainty as a proxy to analyze the bias of 11 models against women and non-whites . |
| Outcome: | The proposed method analyzes the bias of 11 models against women and non-white annotators . it shows that some pre-trained models predict with high accuracy the labels coming from minority groups . |