Papers by Buse Çarık

3 papers
A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions.
Approach: They propose a dataset of Reddit posts annotated across four causal tasks . they use a binary causal classification, explicit vs. implicit causality, cause–effect span extraction and causal gist generation to bridge causal detection and reasoning over informal discourse.
Outcome: The proposed dataset analyzes 10,120 Reddit posts discussing public health related to the COVID-19 pandemic.
A Twitter Corpus for Named Entity Recognition in Turkish (2022.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of information extraction that uses predefined named entities to identify NEs in noisy texts.
Approach: They propose to use a Turkish Twitter Named Entity Recognition dataset to identify predefined named entities (NEs) the dataset contains 5000 tweets from a year-long period with a high agreement score.
Outcome: The proposed dataset contains 5000 tweets from a year-long period and has high agreement scores.
A Turkish Hate Speech Dataset and Detection System (2022.lrec-1)

Copied to clipboard

Challenge: Davidson et al., 2017: hate speech is a discourse that targets a specific group based on race, gender, religion, sexual orientation, etc.
Approach: They propose a machine learning system for automatic detection of hate speech in Turkish . they use a hate speech dataset and a dataset to collect tweets about immigrants .
Outcome: The proposed system is able to detect hate speech in Turkish and annotate it using BERTurk.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations