Papers by Saif Mohammad
WikiArt Emotions: An Annotated Dataset of Emotions Evoked by Art (L18-1)
Copied to clipboard
| Challenge: | a dataset of 4,000 pieces of art has annotations for emotions evoked in the observer . the dataset can help answer questions about what makes art evocative, how does art convey different emotions, what attributes of a painting make it well liked, and how much does the title impact the affectual response to art. |
| Approach: | They create a dataset of 4,000 western art pieces that has annotations for emotions . they use crowdsourcing to annotate the art for one or more of twenty emotion categories . fear, happiness, love, sadness were the dominant emotions that obtained consistent annotations . |
| Outcome: | The dataset shows that the most popular emotions are fear, happiness, love and sadness . the dataset can be used to develop systems that detect emotions evoked by art . |
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)
Copied to clipboard
| Challenge: | a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories . |
| Approach: | They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets . |
| Outcome: | The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets . |
Forgotten Knowledge: Examining the Citational Amnesia in NLP (2023.acl-long)
Copied to clipboard
| Challenge: | a recent study examines how far back in time we tend to cite papers . citation patterns are correlated with age, age, and other factors . |
| Approach: | They analyze citation patterns across time and examine temporal changes . they find that 62% of cited papers are from the immediate five years prior to publication . |
| Outcome: | The authors show that citing papers is the primary method of scientific writing . they show that the trend has reversed and current papers have low temporal diversity . |
Word Affect Intensities (L18-1)
Copied to clipboard
| Challenge: | Existing lexicons of affect only show coarse associations, but are not accurate as human-created ones. |
| Approach: | They propose to use a manually created affect intensity lexicon with real-valued intensity scores for anger, fear, joy, and sadness. |
| Outcome: | The lexicon has real-valued scores for anger, fear, joy, and sadness . anger, fears, and sad words have very similar VAD scores . |
Quantifying Qualitative Data for Understanding Controversial Issues (L18-1)
Copied to clipboard
| Challenge: | 'Controversy' is a state of sustained public debate on a topic or issue that evokes conflicting opinions, beliefs, claims, arguments, and points of view. |
| Approach: | They propose a crowdsourced approach to quantifying qualitative information on controversial issues by analyzing crowdsourced assertions in social media. |
| Outcome: | The proposed dataset consists of over 2,000 assertions on 16 controversial issues. |
WorryWords: Norms of Anxiety Association for over 44k English Words (2024.emnlp-main)
Copied to clipboard
| Challenge: | Anxiety is a common and beneficial human emotion, but there is still much that is not known about it . |
| Approach: | They propose a repository of manually derived word–anxiety associations for over 44,450 English words. |
| Outcome: | The proposed system can track anxiety in streams of text using WorryWords alone. |
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)
Copied to clipboard
Nedjma Ousidhoum, Shamsuddeen Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Ahmad, Sanchit Ahuja, Alham Aji, Vladimir Araujo, Abinew Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine Kock, Genet Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Tilaye, Krishnapriya Vishnubhotla, Genta Winata, Seid Yimam, Saif Mohammad
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
Big BiRD: A Large, Fine-Grained, Bigram Relatedness Dataset for Examining Semantic Composition (N19-1)
Copied to clipboard
| Challenge: | Existing datasets of semantic relatedness only include pairs of unigrams (single words) Existing data suffer from inconsistent annotations and scale region bias due to rating scales. |
| Approach: | They propose to use a large, fine-grained, bigram relatedness dataset to compare the relatedness of 3,345 English term pairs using a comparative annotation technique called Best–Worst Scaling. |
| Outcome: | The proposed datasets are highly reliable and have a split-half reliability of 0.937. |
Emotion Granularity from Text: An Aggregate-Level Indicator of Mental Health (2024.emnlp-main)
Copied to clipboard
| Challenge: | Emotions play a central role in how we construct meaning and communicate with others. |
| Approach: | They propose to use temporally-ordered speaker utterances to measure emotion granularity in social media to determine whether they are effective as mental health markers. |
| Outcome: | The proposed measures of emotion granularity function as markers of mental health conditions. |
D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research (2022.lrec-1)
Copied to clipboard
| Challenge: | DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. |
| Approach: | They extracted metadata from more than 6 million DBLP publications to create the DB3 Discovery Dataset (D3) . they found that computer science is a growing research field (15% annually), with an active and collaborative researcher community. |
| Outcome: | The DBLP Discovery Dataset (D3) can be used to identify trends in research activity, productivity, focus, bias, accessibility, and impact of computer science research. |
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)
Copied to clipboard
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, Stephen Arthur
| Challenge: | Africa has the highest linguistic diversity among all continents. |
| Approach: | They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges . |
| Outcome: | The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task . |
A Diachronic Analysis of Paradigm Shifts in NLP Research: When, How, and Why? (2023.emnlp-main)
Copied to clipboard
| Challenge: | a systematic framework to analyze the evolution of research topics in a scientific field is crucial for keeping abreast of its continuous advancement. |
| Approach: | They propose a framework for analyzing the evolution of research topics in a scientific field using causal discovery and inference techniques. |
| Outcome: | The proposed framework uncovers evolutionary trends and causes for a wide range of NLP topics. |
Evaluating Emotion Arcs Across Languages: Bridging the Global Divide in Sentiment Analysis (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Emotion arcs capture how an individual (or a population) feels over time. |
| Approach: | They compare machine-learning and Lexicon-Only methods to generate emotion arcs . they run experiments on 18 diverse datasets in 9 languages . |
| Outcome: | The proposed method is poor at instance level emotion classification, but highly accurate when aggregating information from hundreds of instances. |
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)
Copied to clipboard
| Challenge: | In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other) |
| Approach: | They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 . |
| Outcome: | The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show . |
Language and Mental Health: Measures of Emotion Dynamics from Text as Linguistic Biosocial Markers (2023.emnlp-main)
Copied to clipboard
| Challenge: | valence variability was significantly lower in the control group compared to ADHD, depression, bipolar disorder, MDD, PTSD, and OCD but not PPD. |
| Approach: | They study the relationship between tweet emotion dynamics and mental health disorders by using a user-disclosed diagnosis. |
| Outcome: | The results show that the measures varied by the user's self-disclosed diagnosis. |
Ethics Sheets for AI Tasks (2022.acl-long)
Copied to clipboard
| Challenge: | a recent study has shown that technology can lead to more adverse outcomes for marginalized populations . a new effort is called Ethics Sheets for AI Tasks to flesh out ethical considerations . |
| Approach: | a new effort will focus on ethical considerations at the level of AI tasks . authors propose a template for ethics sheets with 50 ethical consideration examples . |
| Outcome: | a new form of ethics sheets for AI tasks aims to flesh out assumptions and ethical considerations hidden in how a task is commonly framed . a template for ethics sheets with 50 ethical consideration, using the task of emotion recognition as an example, will be presented . |
The Emotion Dynamics of Literary Novels (2024.findings-acl)
Copied to clipboard
| Challenge: | a new study examines the emotional journeys of characters in novels . previous studies have considered a novel as representing a single story arc . |
| Approach: | They analyze the emotion arcs of English literary novels using Utterance Emotion Dynamics . they find that narration and dialogue largely express disparate emotions through the course of a novel . |
| Outcome: | The analysis of English literary novels shows that narration and dialogue express disparate emotions . the commonalities or differences in the emotional arcs are more accurately captured by individual characters . |
What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing work on semantic relatedness has focused on semantic similarity because of a lack of relatedness datasets. |
| Approach: | They propose a dataset for semantic relatedness that has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
| Outcome: | The proposed dataset has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
Best Practices in the Creation and Use of Emotion Lexicons (2023.findings-eacl)
Copied to clipboard
| Challenge: | Inappropriate and incorrect use of emotion lexicons can lead to harmful inferences . |
| Approach: | They propose to present some of the practical and ethical considerations involved in the creation and use of emotion lexicons. |
| Outcome: | The proposed lexicons can lead to harmful inferences and sub-optimal results . the aim is to provide a comprehensive set of practical and ethical considerations . |
Obtaining Reliable Human Ratings of Valence, Arousal, and Dominance for 20,000 English Words (P18-1)
Copied to clipboard
| Challenge: | Words play a central role in language and thought. |
| Approach: | They propose a Lexicon with ratings of valence, arousal, and dominance for 20,000 words . they use Best–Worst Scaling to obtain fine-grained scores . |
| Outcome: | The proposed Lexicon has human ratings of valence, arousal, and dominance for 20,000 words . the ratings are more reliable than those in existing lexicons, the authors show . |
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research (2023.acl-long)
Copied to clipboard
Mohamed Abdalla, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort
| Challenge: | Recent advances in deep learning methods for natural language processing (NLP) have created new business opportunities and made NLP research critical for industry development. |
| Approach: | They examine industry presence in the field since the early 90s and characterize it using a corpus of 78,187 NLP publications and 701 resumes of NLP publication authors. |
| Outcome: | The authors find that industry presence among NLP authors has been steady before a steep increase over the past five years (180% growth from 2017 to 2022). |