Papers by Buse Çarık
A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions. |
| Approach: | They propose a dataset of Reddit posts annotated across four causal tasks . they use a binary causal classification, explicit vs. implicit causality, cause–effect span extraction and causal gist generation to bridge causal detection and reasoning over informal discourse. |
| Outcome: | The proposed dataset analyzes 10,120 Reddit posts discussing public health related to the COVID-19 pandemic. |
A Twitter Corpus for Named Entity Recognition in Turkish (2022.lrec-1)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a subtask of information extraction that uses predefined named entities to identify NEs in noisy texts. |
| Approach: | They propose to use a Turkish Twitter Named Entity Recognition dataset to identify predefined named entities (NEs) the dataset contains 5000 tweets from a year-long period with a high agreement score. |
| Outcome: | The proposed dataset contains 5000 tweets from a year-long period and has high agreement scores. |
A Turkish Hate Speech Dataset and Detection System (2022.lrec-1)
Copied to clipboard
| Challenge: | Davidson et al., 2017: hate speech is a discourse that targets a specific group based on race, gender, religion, sexual orientation, etc. |
| Approach: | They propose a machine learning system for automatic detection of hate speech in Turkish . they use a hate speech dataset and a dataset to collect tweets about immigrants . |
| Outcome: | The proposed system is able to detect hate speech in Turkish and annotate it using BERTurk. |