Papers by Georgi Karadzhov

6 papers
SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification (2021.findings-acl)

Copied to clipboard

Challenge: toxicity, hate speech, cyberbullying, and cyber-aggression are common themes in social media . authors present a dataset that is limited in size and biased towards offensive language .
Approach: They present an expanded dataset that uses a taxonomy for offensive language identification . they show that using SOLID and OLID yields sizable performance gains .
Outcome: The proposed dataset shows that it performs better than the OLID dataset for two different models.
Predicting Factuality of Reporting and Bias of News Media Sources (D18-1)

Copied to clipboard

Challenge: a new study examines the factuality of news media and its biases . social media has democratized content creation and spread information online .
Approach: They propose to characterize entire news media to predict factuality and bias . they experiment with news websites and a set of features derived from their content .
Outcome: The proposed model shows that the features of news websites perform better than baseline . the results show that the feature types are important for fact-checking systems .
Tanbih: Get To Know What You Are Reading (D19-3)

Copied to clipboard

Challenge: Nowadays, more and more readers consume news online.
Approach: They propose a news platform that displays news grouped into events and generates media profiles that show the general factuality of reporting, the degree of propagandistic content, hyper-partisanship, leading political ideology, general frame of reporting and stance with respect to various claims and topics of a media outlet.
Outcome: The proposed news platform displays news grouped into events and generates media profiles that show the factuality of reporting, the degree of propagandistic content, hyper-partisanship, leading political ideology, general frame of reporting and stance with respect to various claims and topics of a news outlet.
Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models (2025.acl-long)

Copied to clipboard

Challenge: Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text.
Approach: They propose a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance.
Outcome: The proposed framework improves diffusion-based text generation and improves scalability and fluency.
Multi-Task Ordinal Regression for Jointly Predicting the Trustworthiness and the Leading Political Ideology of News Media (N19-1)

Copied to clipboard

Challenge: a number of fact-checking initiatives have been launched, both manual and automatic, but the whole enterprise remains in a state of crisis.
Approach: They propose a multi-task ordinal regression framework that models trustworthiness estimation and political ideology detection of entire news outlets.
Outcome: The proposed model outperforms models that target the problems in isolation.
What Was Written vs. Who Read It: News Media Profiling Using Text Analysis and Social Media Context (2020.acl-main)

Copied to clipboard

Challenge: a growing number of fake news reports are published online, causing a trust crisis . a new study aims to predict political bias and factuality of reporting of entire news outlets .
Approach: They propose to profile entire news outlets and look for those that are likely to publish fake content . they also examine what was written about the target medium and who reads it .
Outcome: The proposed method improves on the current state-of-the-art in analyzing social media and what was written about the target medium.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations