Papers by Cagri Toraman
D2U: Distance-to-Uniform Learning for Out-of-Scope Detection (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for out-of-scope (OOS) detection use classifier confidence score, but model cannot infer correctly. |
| Approach: | They propose a zero-shot post-processing step that exploits the classification confidence score and the shape of the entire output distribution. |
| Outcome: | The proposed method improves performance when there is no OOS training data and learning procedure when OOS data is available. |
JL-Hate: An Annotated Dataset for Joint Learning of Hate Speech and Target Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing data resources for the detection of hate speech focus on text sequence classification, but the target of hateful content is lacking. |
| Approach: | They propose a tweet dataset for the task of joint learning of hate speech detection and target detection called JL-Hate. |
| Outcome: | The proposed dataset performs similar tasks to the existing datasets in sequence and token classification tasks. |
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | a new dataset of misinformation labels is being developed to detect misinformation on social media platforms . misinformation is spread in many domains including but not limited to health, politics, and disasters . |
| Approach: | They construct a dataset of 5,284 English and 5,064 Turkish tweets with misinformation labels . they use the dataset to analyze misinformation spread and to evaluate misinformation detection . |
| Outcome: | The proposed dataset includes 5,284 English and 5,064 Turkish tweets with misinformation labels for several recent events between 2020 and 2022. |
Large-Scale Hate Speech Detection with Cross-Domain Transfer (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets for hate speech detection are limited due to the labor cost. |
| Approach: | They construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each. |
| Outcome: | The proposed datasets outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection. |
PejorativITy: Disambiguating Pejorative Epithets to Improve Misogyny Detection in Italian Tweets (2024.lrec-main)
Copied to clipboard
Arianna Muti, Federico Ruggeri, Cagri Toraman, Alberto Barrón-Cedeño, Samuel Algherini, Lorenzo Musetti, Silvia Ronchi, Gianmarco Saretto, Caterina Zapparoli
| Challenge: | Disambiguating the meaning of pejorative words might help misogyny detection . state-of-the-art models struggle to correctly classify misogoyne when sentences contain such terms. |
| Approach: | They present a corpus of 1,200 manually annotated Italian tweets for pejorative language at the word level and misogyny at the sentence level. |
| Outcome: | The proposed model improves on 1,200 manually annotated Italian tweets and on two benchmarks. |