The PRECOM-SM Corpus: Gambling in Spanish Social Media (2025.coling-main)

Copied to clipboard

Challenge: ESTUDES de Sanidad: 20.1% of youngsters between 14 and 18 years old have gambled money online . ESTUDES: 2021: 17.9% of students who have gamble would be predisposed to gambling-related problems.
Approach: This paper collects text from online Spanish-speaking communities and analyses it to detect gambling addiction problem.
Outcome: The proposed corpus collects text from Spanish-speaking communities and analyzes it . it finds patterns in written language from frequent and infrequent users . 20.1% of youngsters between 14 and 18 years old have gambled money in person or online .

Similar Papers

Analyzing Gambling Addictions: A Spanish Corpus for Understanding Pathological Behavior (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the interaction between natural language use and gambling disorders.
Approach: They build a new corpus of sentences that are searched and compared using top-k pooling to form the assessment pools of sentences.
Outcome: The proposed model is based on a new corpus of sentences in spanish .
MentalRiskES: A New Corpus for Early Detection of Mental Disorders in Spanish (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on the prevalence of mental disorders on the Web are limited to the English language.
Approach: They propose to use user messages posted on Telegram groups to annotate the corpus for natural language processing and to conduct experiments on text classification and regression.
Outcome: The proposed corpus contains over 1,300 subjects with more than 45,000 messages posted in different public Telegram groups.
Abusive language in Spanish children and young teenager’s conversations: data preparation and short text classification with contextual word embeddings (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on how to automatically detect abusive short texts are gaining interest in the natural language processing community.
Approach: They propose to use a contextual word embedding model to automatically detect abusive short texts for Spanish language.
Outcome: The proposed model outperforms classical methods in the detection of abusive short texts for the spanish language.
Searching Brazilian Twitter for Signs of Mental Health Issues (2020.lrec-1)

Copied to clipboard

Challenge: Existing resources are largely devoted to English NLP, and there is little support for these studies in under resourced languages.
Approach: They propose to build a corpus in Brazilian Portuguese to support both the recognition of mental health issues and the temporal analysis of these illnesses.
Outcome: The proposed corpus will support both the recognition of mental health issues and the temporal analysis of these illnesses in the Brazilian Portuguese language.
RoBERTuito: a pre-trained language model for social media text in Spanish (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained language models have been used in many natural language processing tasks . some domain-specific models have shown to improve performance in some domains . however, for languages other than English, such models are not widely available .
Approach: They present a pre-trained language model for user-generated text in Spanish . it is based on 500 million tweets and has some cross-lingual abilities .
Outcome: The model outperforms models trained on over 500 million tweets on a benchmark in spanish and english.
A Dataset for Multi-lingual Epidemiological Event Extraction (2020.lrec-1)

Copied to clipboard

Challenge: Using the Web, we propose a corpus for information extraction and text classification.
Approach: They propose to use a corpus for information extraction and natural language processing (NLP) tasks such as text classification.
Outcome: The proposed corpus can be used for information extraction and natural language processing tasks such as text classification.
HAHA 2019 Dataset: A Corpus for Humor Analysis in Spanish (2020.lrec-1)

Copied to clipboard

Challenge: 30,000 Spanish tweets were crowd-annotated with humor value and funniness score . the corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweet.
Approach: They develop a corpus of 30,000 Spanish tweets crowd-annotated with humor value and funniness score.
Outcome: The results obtained from the 30,000 tweets in the Spanish language are encouraging.
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting social-media texts are limited to the English language and longer texts are not easily recognisable by humans.
Approach: They propose to use a multilingual and multi-platform dataset to compare machine-generated text detection methods in the social-media domain to compare them to human-written texts.
Outcome: The proposed dataset contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs.
Words are the Window to the Soul: Language-based User Representations for Fake News Detection (2020.coling-main)

Copied to clipboard

Challenge: Existing studies on fake news classification focus on textual content, but also social context in which news are consumed.
Approach: They propose a model that creates representations of individuals on social media based only on the language they produce and uses them to detect fake news.
Outcome: The proposed model exploits the relationship between language use and connections in the social graph to assess the presence of the Echo Chamber effect in the data.
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats (2025.acl-long)

Copied to clipboard

Challenge: Dog whistles are coded expressions with dual meanings that slip by content moderation filters . a new study finds that state-of-the-art systems fail to identify novel dog whistles .
Approach: They propose a task to find novel dog whistles in massive social media corpora . they use a strong baseline system that combines vector databases and Large Language Models to identify new dog whistle.
Outcome: The proposed system fails to identify dog whistles across three social media cases . it combines vector databases and Large Language Models to efficiently and effectively identify new dog whistle expressions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations