Challenge: a new study aims to detect propaganda in multiple languages using code-switching . social media platforms have made it easier for anyone to spread information to a wide audience .
Approach: They propose to detect propaganda techniques in code-switched texts using a corpus of 1,030 texts . they propose to model multilinguality directly rather than using translation .
Outcome: The proposed method combines different languages within the same text, presenting a challenge for automatic systems.

Similar Papers

Offensive Content Detection via Synthetic Code-Switched Text (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to detect offensive content in social media platforms are limited by the availability of labeled code-switched data.
Approach: They propose a method for generating synthetic code-switched offensive content data using human-generated data and a keyword classification baseline.
Outcome: The proposed algorithm can be used to generate synthetic code-switched offensive content data and train it on human-generated data.
Detecting Propaganda Techniques in Memes (2021.acl-long)

Copied to clipboard

Challenge: Propaganda can be defined as a form of communication that aims to influence opinions or the actions of people towards a specific goal.
Approach: They propose to detect the type of propaganda techniques used in memes by annotating them with 22 techniques.
Outcome: The proposed model identifies 22 propaganda techniques in memes, which can appear in text, image or both .
TWEETSPIN: Fine-grained Propaganda Detection in Social Media Using Multi-View Representations (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies on propaganda detection involve document and fragment-level analyses of news articles.
Approach: They propose a neural approach to detect and categorize propaganda tweets across fine-grained categories . they use a dataset containing tweets weakly annotated with different propaganda techniques .
Outcome: The proposed method outperforms benchmark methods and transfers knowledge to low-resource news domains.
Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Using large language models (LLMs) to detect propaganda from text is a challenge for the development of sophisticated models.
Approach: They propose to use a large propaganda dataset to identify propagandistic content in text, visual, or multimodal languages to improve their models.
Outcome: The proposed model performs better on a large propaganda dataset than the existing models on skewed datasets.
Collecting Code-Switched Data from Social Media (L18-1)

Copied to clipboard

Challenge: a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages .
Approach: They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets .
Outcome: The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets .
Improved Sentiment Detection via Label Transfer from Monolingual to Synthetic Code-Switched Text (P19-1)

Copied to clipboard

Challenge: Existing sentiment detection methods are trained on sentiment-labeled monolingual text.
Approach: They propose a method for synthesizing labeled code-switched text from monolingual text.
Outcome: The proposed method improves sentiment labeling accuracy for three languages.
Unleashing the Power of Discourse-Enhanced Transformers for Propaganda Detection (2024.eacl-long)

Copied to clipboard

Challenge: Existing systems focused on the surface words, ignoring the linguistic structure of the texts.
Approach: They propose to use discourse analysis to analyze paragraph-level and token-level classifications and propose a Transformer architecture that can be used to detect propaganda.
Outcome: The proposed system improves on English and Russian texts and shows strong correlations between propaganda instances and discourse spans.
Code-Mixed Probes Show How Pre-Trained Models Generalise on Code-Switched Text (2024.lrec-main)

Copied to clipboard

Challenge: Code-switching is a prevalent linguistic phenomenon in which multilingual individuals seamlessly alternate between languages.
Approach: They propose to use pre-trained language models to generalise to code-switched text . they use a dataset of well-formed naturalistic code-witched texts and parallel translations into the source languages to examine their results.
Outcome: The proposed model generalises to code-switched text, shedding light on their ability to generalise representations to CS corpora.
Synthetic Propaganda Embeddings To Train A Linear Projection (D19-50)

Copied to clipboard

Challenge: Using contextualized token embeddings, we can extract features of propaganda from contextualized embeddnings without fine-tuning the large parameters of the base model.
Approach: They propose a method for detecting fine-grained categories of propaganda in text by generating synthetically generated embeddings from pre-trained language models.
Outcome: The proposed method is used in the first shared task in fine-grained propaganda detection at NLP4IF as Team Stalin.
Towards Code-switched Classification Exploiting Constituent Language Resources (2020.aacl-srw)

Copied to clipboard

Challenge: Code-switching is a communicative phenomenon denoting a shift from one language to another within the same speech exchange.
Approach: They propose to convert code-switched data into its constituent high resource languages for use in both monolingual and cross-lingual settings.
Outcome: The proposed code-switching language can be used for multiple downstream tasks . the proposed language increases the F1 score by 22% and 42.5% compared to the state-of-the-art.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations