Challenge: Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing.
Approach: They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets.
Outcome: The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter.

Similar Papers

XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond (2022.lrec-1)

Copied to clipboard

Challenge: Language models are ubiquitous in NLP, but current analyses focus on (multilingual variants of) standard benchmarks and task-specific corpora as multilingual signals.
Approach: They propose a model to train and evaluate multilingual language models in Twitter using a set of Twitter datasets in eight different languages and a XLM-T model.
Outcome: The proposed model trains and evaluates multilingual models on Twitter.
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Several studies investigating methods to detect offensive content in social media use English data.
Approach: They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources.
Outcome: The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish.
Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation (2020.coling-main)

Copied to clipboard

Challenge: Existing models for sentiment analysis over tweets require a substantial amount of text to adapt to a domain where the syntax is different.
Approach: They propose to use a multilingual transformer model to train over tweets in five different languages to adapt the model to non-English languages.
Outcome: The proposed model improves over small corpora of tweets in non-English languages.
Bernice: A Multilingual Pre-trained Encoder for Twitter (2022.emnlp-main)

Copied to clipboard

Challenge: Existing language models for Twitter are monolingual, adapted from other domains, or trained on limited amount of in-domain data.
Approach: They propose a multilingual RoBERTa language model that is trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer.
Outcome: The proposed model outperforms or matches models trained on monolingual and multilingual tweets on a variety of benchmarks and is more efficient compute- and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer.
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data.
Approach: They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria.
Outcome: The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets.
Processing and Understanding Mixed Language Data (D19-2)

Copied to clipboard

Challenge: Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text .
Approach: a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities.
Outcome: a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing.
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)

Copied to clipboard

Challenge: a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages.
Approach: They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction .
Outcome: The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms.
TweetNLP: Cutting-Edge Natural Language Processing for Social Media (2022.emnlp-demos)

Copied to clipboard

Challenge: TweetNLP is an integrated platform for natural language processing in social media.
Approach: They propose a Python-based platform for natural language processing in social media that supports a variety of NLP tasks.
Outcome: The proposed platform supports generic focus areas such as sentiment analysis and named entity recognition, as well as social media-specific tasks such as emoji prediction and offensive language identification.
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on sentiment analysis of tweets focus on the English language . however, there is still a challenge of processing lower-resourced languages .
Approach: They transform tweet sentiment dataset into a multimodal format through a straightforward curation process.
Outcome: The proposed approach performs exceptionally well in unimodal and multimodal configurations.
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting social-media texts are limited to the English language and longer texts are not easily recognisable by humans.
Approach: They propose to use a multilingual and multi-platform dataset to compare machine-generated text detection methods in the social-media domain to compare them to human-written texts.
Outcome: The proposed dataset contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations