Challenge: Existing methods to detect offensive content in social media platforms are limited by the availability of labeled code-switched data.
Approach: They propose a method for generating synthetic code-switched offensive content data using human-generated data and a keyword classification baseline.
Outcome: The proposed algorithm can be used to generate synthetic code-switched offensive content data and train it on human-generated data.

Similar Papers

Word-Level Detection of Code-Mixed Hate Speech with Multilingual Domain Transfer (2025.findings-acl)

Copied to clipboard

Challenge: a growing problem in language detection tasks is code-mixing, a combination of more than one language . lack of available datasets for code-mixing causes the problem . authors propose a multilingual approach to code-matching .
Approach: They propose to use an annotated hate speech dataset to detect code-mixing in profane language . they propose to apply bilingual fine-tuned models to code-mixed hate speech in german rap lyrics .
Outcome: The proposed model can detect code-mixed hate speech and neologisms in German rap lyrics . the proposed model is more nuanced than binary classification .
Offensive Language and Hate Speech Detection for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a growing number of social media platforms are detecting and dealing with offensive language . a recent study found that the best performing system for English is best for Danish .
Approach: They propose automatic methods to detect offensive language on social media platforms . they use user-generated comments from various social media sites to find offensive language .
Outcome: The proposed system performs best for both English and Danish language . it achieves a macro averaged F1-score of 0.74 and a best for Danish achieves 0.73 .
Fighting Offensive Language on Social Media with Unsupervised Text Style Transfer (P18-2)

Copied to clipboard

Challenge: Existing methods to tackle the problem of offensive language in social media are based on machine learning.
Approach: They propose a method for training encoder-decoders using non-parallel data . they use a collaborative classifier, attention and the cycle consistency loss .
Outcome: The proposed method outperforms state-of-the-art text style transfer systems on Twitter and Reddit . it produces reliable non-offensive transferred sentences, the authors show .
Detecting Propaganda Techniques in Code-Switched Social Media Text (2023.emnlp-main)

Copied to clipboard

Challenge: a new study aims to detect propaganda in multiple languages using code-switching . social media platforms have made it easier for anyone to spread information to a wide audience .
Approach: They propose to detect propaganda techniques in code-switched texts using a corpus of 1,030 texts . they propose to model multilinguality directly rather than using translation .
Outcome: The proposed method combines different languages within the same text, presenting a challenge for automatic systems.
Towards Code-switched Classification Exploiting Constituent Language Resources (2020.aacl-srw)

Copied to clipboard

Challenge: Code-switching is a communicative phenomenon denoting a shift from one language to another within the same speech exchange.
Approach: They propose to convert code-switched data into its constituent high resource languages for use in both monolingual and cross-lingual settings.
Outcome: The proposed code-switching language can be used for multiple downstream tasks . the proposed language increases the F1 score by 22% and 42.5% compared to the state-of-the-art.
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Several studies investigating methods to detect offensive content in social media use English data.
Approach: They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources.
Outcome: The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish.
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms.
Approach: They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model.
Outcome: The proposed model improves performance in cross-dataset training.
Offensive Video Detection: Dataset and Baseline Results (2020.lrec-1)

Copied to clipboard

Challenge: a large number of social media platforms discourage users from publishing offensive content . however, there is no method to detect offensive content on these platforms due to the high volume of publications.
Approach: They propose to use text-based machine learning to detect offensive content on different platforms . they use word embedding with Deep Learning classifiers to perform best results .
Outcome: The proposed methods outperform Classic and Deep Learning classifiers in Portuguese and CNN architectures in other features.
Thesis Proposal: An Explainable Multimodal Framework for Detecting Harmful Content in Code-Switched Children’s Media (2026.acl-srw)

Copied to clipboard

Challenge: Current content moderation systems fail to protect children from harmful content, especially in under-resourced, code-switched settings.
Approach: They propose to integrate a fine-tuned classifier with an LLM-powered module that synthesizes the classifier’s internal evidential signals to generate faithful, human-readable rationales for each decision.
Outcome: The proposed framework integrates a fine-tuned classifier for accurate, scalable detection with an LLM-powered module that synthesizes the classifier’s internal evidential signals to generate faithful, human-readable rationales for each decision.
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)

Copied to clipboard

Challenge: Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us .
Approach: They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages.
Outcome: The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations