What does the language of foods say about us? (D19-62)

Copied to clipboard

Challenge: Using a dataset of 24 million food-related tweets, we can predict if states in the United States are above the median rates for type 2 diabetes mellitus (T2DM) income, poverty, and education are important factors in predicting T2DM rates, but socioeconomic factors do not capture this information.
Approach: They use a dataset of 24 million food-related tweets to investigate the signal contained in the language of food on social media.
Outcome: The language of food can predict health risks, political orientation, and geographic location, and outperform previous work by 4–18%.

Similar Papers

Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet lack a robust methodology to dissect these phenomena comprehensively.
Approach: They propose a multilingual dataset centered on food-related cultural facts and variations in food practices.
Outcome: The proposed model incorporates cultural context significantly and improves its ability to access cultural knowledge.
Native Language Identification with User Generated Content (D18-1)

Copied to clipboard

Challenge: Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task.
Approach: They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task.
Outcome: The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers.
Mining Cross-Cultural Differences and Similarities in Social Media (P18-1)

Copied to clipboard

Challenge: a new paper examines the problem of computing cross-cultural differences and similarities in natural language understanding . cross-culture differences are important for cross-lingual research, especially in social media .
Approach: They propose a framework for computing cross-cultural differences and similarities from social media . they propose to use a social media platform to find similar terms for slang across languages .
Outcome: The proposed framework outperforms baseline methods on two novel tasks.
Diachronic degradation of language models: Insights from social media (P18-2)

Copied to clipboard

Challenge: Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language.
Approach: They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling.
Outcome: The results show that it is possible to measure diachronic drifts within social media and within the span of a few years.
FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture (2024.emnlp-main)

Copied to clipboard

Challenge: FoodieQA is a manually curated, fine-grained image-text dataset capturing the intricate features of food cultures across various regions in China.
Approach: They evaluate vision–language Models and large language models on unseen food images and corresponding questions.
Outcome: The proposed dataset evaluates vision–language Models and large language models on unseen food images and corresponding questions.
Measuring Social Biases in Grounded Vision and Language Embeddings (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to measure social biases in word embeddings are limited to visually grounded word embeds . a new study generalizes word embedment associations to visually ground word embeddas .
Approach: They generalize word embeddings' biases to visually grounded word embeds . they propose two generalizations that answer questions about how biase, language, and vision interact .
Outcome: The proposed measures are applied to a new dataset that includes 10,228 images from COCO, Conceptual Captions, and Google Images.
Social Meme-ing: Measuring Linguistic Variation in Memes (2024.naacl-long)

Copied to clipboard

Challenge: In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech.
Approach: They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset.
Outcome: The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes.
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)

Copied to clipboard

Challenge: Vulgarity is a common linguistic expression and is used to perform several linguistic functions.
Approach: They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance.
Outcome: The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings.
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.
Splits! Flexible Sociocultural Linguistic Investigation at Scale (2026.acl-long)

Copied to clipboard

Challenge: Variation in language use offers a rich lens into cultural perspectives, values, and opinions.
Approach: They propose to construct a "sandbox" for systematic and flexible sociolinguistic research by splitting a reddit dataset into demographically/topically split SLPs.
Outcome: The proposed method analyzes a demographically/topically split Reddit dataset validated by self-identification and replicating several known SLPs from existing literature.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations