| Challenge: | Using a dataset of 24 million food-related tweets, we can predict if states in the United States are above the median rates for type 2 diabetes mellitus (T2DM) income, poverty, and education are important factors in predicting T2DM rates, but socioeconomic factors do not capture this information. |
| Approach: | They use a dataset of 24 million food-related tweets to investigate the signal contained in the language of food on social media. |
| Outcome: | The language of food can predict health risks, political orientation, and geographic location, and outperform previous work by 4–18%. |
Similar Papers
Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge (2025.naacl-long)
Copied to clipboard
Li Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, Daniel Hershcovich
| Challenge: | Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet lack a robust methodology to dissect these phenomena comprehensively. |
| Approach: | They propose a multilingual dataset centered on food-related cultural facts and variations in food practices. |
| Outcome: | The proposed model incorporates cultural context significantly and improves its ability to access cultural knowledge. |
Native Language Identification with User Generated Content (D18-1)
Copied to clipboard
| Challenge: | Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task. |
| Approach: | They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task. |
| Outcome: | The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers. |
Mining Cross-Cultural Differences and Similarities in Social Media (P18-1)
Copied to clipboard
| Challenge: | a new paper examines the problem of computing cross-cultural differences and similarities in natural language understanding . cross-culture differences are important for cross-lingual research, especially in social media . |
| Approach: | They propose a framework for computing cross-cultural differences and similarities from social media . they propose to use a social media platform to find similar terms for slang across languages . |
| Outcome: | The proposed framework outperforms baseline methods on two novel tasks. |
Diachronic degradation of language models: Insights from social media (P18-2)
Copied to clipboard
| Challenge: | Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language. |
| Approach: | They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling. |
| Outcome: | The results show that it is possible to measure diachronic drifts within social media and within the span of a few years. |
FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture (2024.emnlp-main)
Copied to clipboard
Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott
| Challenge: | FoodieQA is a manually curated, fine-grained image-text dataset capturing the intricate features of food cultures across various regions in China. |
| Approach: | They evaluate vision–language Models and large language models on unseen food images and corresponding questions. |
| Outcome: | The proposed dataset evaluates vision–language Models and large language models on unseen food images and corresponding questions. |
Measuring Social Biases in Grounded Vision and Language Embeddings (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to measure social biases in word embeddings are limited to visually grounded word embeds . a new study generalizes word embedment associations to visually ground word embeddas . |
| Approach: | They generalize word embeddings' biases to visually grounded word embeds . they propose two generalizations that answer questions about how biase, language, and vision interact . |
| Outcome: | The proposed measures are applied to a new dataset that includes 10,228 images from COCO, Conceptual Captions, and Google Images. |
Social Meme-ing: Measuring Linguistic Variation in Memes (2024.naacl-long)
Copied to clipboard
| Challenge: | In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech. |
| Approach: | They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset. |
| Outcome: | The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes. |
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)
Copied to clipboard
| Challenge: | Vulgarity is a common linguistic expression and is used to perform several linguistic functions. |
| Approach: | They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance. |
| Outcome: | The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings. |
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature . |
| Approach: | They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language. |
| Outcome: | The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders. |
Splits! Flexible Sociocultural Linguistic Investigation at Scale (2026.acl-long)
Copied to clipboard
| Challenge: | Variation in language use offers a rich lens into cultural perspectives, values, and opinions. |
| Approach: | They propose to construct a "sandbox" for systematic and flexible sociolinguistic research by splitting a reddit dataset into demographically/topically split SLPs. |
| Outcome: | The proposed method analyzes a demographically/topically split Reddit dataset validated by self-identification and replicating several known SLPs from existing literature. |