| Challenge: | Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. |
| Approach: | They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets. |
| Outcome: | The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter. |
Similar Papers
XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond (2022.lrec-1)
Copied to clipboard
| Challenge: | Language models are ubiquitous in NLP, but current analyses focus on (multilingual variants of) standard benchmarks and task-specific corpora as multilingual signals. |
| Approach: | They propose a model to train and evaluate multilingual language models in Twitter using a set of Twitter datasets in eight different languages and a XLM-T model. |
| Outcome: | The proposed model trains and evaluates multilingual models on Twitter. |
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Several studies investigating methods to detect offensive content in social media use English data. |
| Approach: | They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources. |
| Outcome: | The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish. |
Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing models for sentiment analysis over tweets require a substantial amount of text to adapt to a domain where the syntax is different. |
| Approach: | They propose to use a multilingual transformer model to train over tweets in five different languages to adapt the model to non-English languages. |
| Outcome: | The proposed model improves over small corpora of tweets in non-English languages. |
Bernice: A Multilingual Pre-trained Encoder for Twitter (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing language models for Twitter are monolingual, adapted from other domains, or trained on limited amount of in-domain data. |
| Approach: | They propose a multilingual RoBERTa language model that is trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer. |
| Outcome: | The proposed model outperforms or matches models trained on monolingual and multilingual tweets on a variety of benchmarks and is more efficient compute- and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer. |
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)
Copied to clipboard
Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, Pavel Brazdil
| Challenge: | Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. |
| Approach: | They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria. |
| Outcome: | The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets. |
Processing and Understanding Mixed Language Data (D19-2)
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
TweetNLP: Cutting-Edge Natural Language Processing for Social Media (2022.emnlp-demos)
Copied to clipboard
Jose Camacho-collados, Kiamehr Rezaee, Talayeh Riahi, Asahi Ushio, Daniel Loureiro, Dimosthenis Antypas, Joanne Boisson, Luis Espinosa Anke, Fangyu Liu, Eugenio Martínez Cámara
| Challenge: | TweetNLP is an integrated platform for natural language processing in social media. |
| Approach: | They propose a Python-based platform for natural language processing in social media that supports a variety of NLP tasks. |
| Outcome: | The proposed platform supports generic focus areas such as sentiment analysis and named entity recognition, as well as social media-specific tasks such as emoji prediction and offensive language identification. |
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on sentiment analysis of tweets focus on the English language . however, there is still a challenge of processing lower-resourced languages . |
| Approach: | They transform tweet sentiment dataset into a multimodal format through a straightforward curation process. |
| Outcome: | The proposed approach performs exceptionally well in unimodal and multimodal configurations. |
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting social-media texts are limited to the English language and longer texts are not easily recognisable by humans. |
| Approach: | They propose to use a multilingual and multi-platform dataset to compare machine-generated text detection methods in the social-media domain to compare them to human-written texts. |
| Outcome: | The proposed dataset contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. |