Papers by Sophie Wu

6 papers
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)

Copied to clipboard

Challenge: Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages.
Approach: They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets.
Outcome: The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers.
Efficient Annotator Reliability Assessment and Sample Weighting for Knowledge-Based Misinformation Detection on Social Media (2025.findings-naacl)

Copied to clipboard

Challenge: Misinformation spreads rapidly on social media, confusing the truth and targeting potentially vulnerable people.
Approach: They propose to use inter- and intra-annotator agreement to understand the reliability of each annotator and influence the training of large language models based on annotators reliability.
Outcome: The proposed framework utilises inter- and intra-annotator agreement to understand the reliability of each annotator and influence the training of large language models based on annotators reliability.
Homophone2Vec: Embedding Space Analysis for Empirical Evaluation of Phonological and Semantic Similarity (2024.acl-srw)

Copied to clipboard

Challenge: Existing studies have shown that homophones with different semantic/syntactic contexts are easier for children to memorize.
Approach: They propose a method for empirically evaluating the relationship between phonological and semantic similarity of linguistic units using embedding spaces.
Outcome: The proposed method shows that Chinese character homophones have a positive semantic relationship at varying levels of sound-sharing.
Probing Narrative Morals: A New Character-Focused MFT Framework for Use with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to categorize moral foundations in storytelling are limited.
Approach: They propose a character-centric method to quantify moral foundations in storytelling using large language models and a novel Moral Foundations Character Action Questionnaire to validate their approach against human annotations.
Outcome: The proposed method validates against human annotations and then applies to 2,697 folktales from 55 countries.
Confabulation: The Surprising Value of Large Language Model Hallucinations (2024.acl-long)

Copied to clipboard

Challenge: 'confabulations' are inherently problematic and AI research should eliminate this flaw, but confabulation is not a problem.
Approach: They argue that measurable semantic characteristics of large language model (LLM) hallucinations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication.
Outcome: The proposed study shows that measurable semantic characteristics of LLM confabulations mirror human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication.
The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions (2025.findings-emnlp)

Copied to clipboard

Challenge: 5% to 10% of posts include body part mentions in English text . text containing BPMs tends to be more emotionally charged, even when the BPM is not used to describe a physical reaction to the emotion in the text.
Approach: They create corpora of body part mentions in online English text with human annotations for the emotions of the person whose body part is mentioned.
Outcome: The proposed study is the first to investigate the connection between emotion, embodiment, and everyday language in a large sample of natural language data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations