Papers by Cagri Toraman

5 papers
D2U: Distance-to-Uniform Learning for Out-of-Scope Detection (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for out-of-scope (OOS) detection use classifier confidence score, but model cannot infer correctly.
Approach: They propose a zero-shot post-processing step that exploits the classification confidence score and the shape of the entire output distribution.
Outcome: The proposed method improves performance when there is no OOS training data and learning procedure when OOS data is available.
JL-Hate: An Annotated Dataset for Joint Learning of Hate Speech and Target Detection (2024.lrec-main)

Copied to clipboard

Challenge: Existing data resources for the detection of hate speech focus on text sequence classification, but the target of hateful content is lacking.
Approach: They propose a tweet dataset for the task of joint learning of hate speech detection and target detection called JL-Hate.
Outcome: The proposed dataset performs similar tasks to the existing datasets in sequence and token classification tasks.
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset of misinformation labels is being developed to detect misinformation on social media platforms . misinformation is spread in many domains including but not limited to health, politics, and disasters .
Approach: They construct a dataset of 5,284 English and 5,064 Turkish tweets with misinformation labels . they use the dataset to analyze misinformation spread and to evaluate misinformation detection .
Outcome: The proposed dataset includes 5,284 English and 5,064 Turkish tweets with misinformation labels for several recent events between 2020 and 2022.
Large-Scale Hate Speech Detection with Cross-Domain Transfer (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for hate speech detection are limited due to the labor cost.
Approach: They construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each.
Outcome: The proposed datasets outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection.
PejorativITy: Disambiguating Pejorative Epithets to Improve Misogyny Detection in Italian Tweets (2024.lrec-main)

Copied to clipboard

Challenge: Disambiguating the meaning of pejorative words might help misogyny detection . state-of-the-art models struggle to correctly classify misogoyne when sentences contain such terms.
Approach: They present a corpus of 1,200 manually annotated Italian tweets for pejorative language at the word level and misogyny at the sentence level.
Outcome: The proposed model improves on 1,200 manually annotated Italian tweets and on two benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations