Crowdsourcing a Large Corpus of Clickbait on Twitter (C18-1)

Copied to clipboard

Challenge: Clickbait is a nuisance on social media.
Approach: a corpus of 38,517 annotated Twitter tweets was constructed to detect clickbait . the corpus was annotating tweets on 4-point scale by five annotators at Amazon's Mechanical Turk .
Outcome: The corpus of 38,517 annotated Twitter tweets was used to evaluate 12 clickbait detectors submitted to the Clickbait Challenge 2017 .

Similar Papers

A French Corpus for Event Detection on Twitter (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets may have different definitions of event or topic, which leads to inconsistent results.
Approach: They present a corpus annotated for event detection tasks consisting of 38 million tweets in French and 130,000 manually annotating tweets as related or unrelated to a given event.
Outcome: The proposed method performs best on 38 million tweets in French and another publicly available dataset of tweets.
An Annotated Corpus for Sexism Detection in French Tweets (2020.lrec-1)

Copied to clipboard

Challenge: Social media networks allow users to share opinions and sentiments, which can cause a large spreading of hatred or abusive messages.
Approach: They propose to annotate 12,000 tweets with a sexism detection scheme in France . they propose to use deep learning to detect if a message with sexist content is really s.
Outcome: The proposed scheme detects sexist content and identifies if it is really sexism . the proposed scheme is the first of its kind in the u.s.
Know Better – A Clickbait Resolving Challenge (2022.lrec-1)

Copied to clipboard

Challenge: a clickbait headline or teaser is used to "bait" the reader into clicking a link to an article . clickbaiting is annoying but effective, and can be countered with specialized models .
Approach: They propose to construct approaches that can automatically extract relevant information from clickbait articles . they argue that clickbaiting can probably not be defeated with clickbaitting detection alone .
Outcome: The proposed methods outperform question answering models on clickbait resolving task . the data will be used to give users tools to counter clickbaiting in the future .
Clickbait Spoiling via Question Answering and Passage Retrieval (2022.acl-long)

Copied to clipboard

Challenge: Clickbait is a term used to describe posts intended to entice readers to visit a web page . clickbait spoiling is generating a short text that satisfies the curiosity induced by a clickbaiting post .
Approach: They propose to use clickbait spoiling to generate a short text that satisfies curiosity . they classify the type of spoiler needed and generate appropriate spoilers .
Outcome: The proposed method outperforms all other methods in generating spoilers for both types of clickbait posts.
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions.
Approach: They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language .
Outcome: The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language .
Annotating the Tweebank Corpus on Named Entity Recognition and Building NLP Models for Social Media Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media data such as Twitter messages pose a particular challenge to NLP systems because of their short, noisy nature.
Approach: They create a Twitter-based NER corpus and train Tweet NLP models on it . they annotate named entities in TB2 using Amazon Mechanical Turk .
Outcome: The proposed model outperforms existing models on Twitter and other social media platforms.
A Novel Contrastive Learning Method for Clickbait Detection on RoCliCo: A Romanian Clickbait Corpus of News Articles (2023.findings-emnlp)

Copied to clipboard

Challenge: Clickbait detection is a task that aims to automatically detect misleading news titles . despite the importance of the task, there is no publicly available clickbait corpus for Romanian .
Approach: They propose a Romanian Clickbait Corpus that automatically detects misleading news titles . they propose four machine learning methods to establish competitive baselines .
Outcome: The proposed model can learn to encode news titles and contents into a deep metric space . the proposed model is available for download on github.com/dariabroscoteanu/RoCliCo.
TWEETQA: A Social Media Focused Question Answering Dataset (P19-1)

Copied to clipboard

Challenge: Social media is becoming an important realtime information source, especially during natural disasters and emergencies.
Approach: They present a large-scale dataset for question answering over social media data . they gather tweets used by journalists and ask human annotators to write questions upon them .
Outcome: The proposed dataset shows that neural models that perform well on formal texts are limited in their performance . the proposed model is still lagging behind human performance with a large margin .
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations