Papers by Sebastiano Salto
Delving into Qualitative Implications of Synthetic Data for Hate Speech Detection (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent work on synthetic data for training models for NLP tasks reports mixed results on subjective tasks such as hate speech detection. |
| Approach: | They propose to use synthetic data to train models for highly subjective tasks such as hate speech detection to investigate the potential and specific pitfalls of using it. |
| Outcome: | The proposed model outperforms models trained with real data on hate speech detection tasks, but it fails to accurately reflect real-world data on linguistic dimensions and results in different class distributions. |