Challenge: a new dataset is being developed to categorize posts that show distress or urgency . the dataset could improve humanitarian efforts, allowing for quicker and more targeted help .
Approach: They propose a dataset that brings together social media posts in the Ukrainian language for the detection of help-seeking posts in times of war.
Outcome: The proposed dataset can be used to improve humanitarian efforts . it can be compared with existing datasets and achieve an accuracy of 81.15% .

Similar Papers

M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signals (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for detecting mental health disorders do not identify individuals actively seeking help.
Approach: This paper introduces a new social media dataset specifically designed to detect help-seeking behavior on social media.
Outcome: The proposed dataset can detect help-seeking behavior on social media . it can address three key tasks: identifying help- seekkers, diagnosing mental health conditions .
Challenges and Opportunities in Information Manipulation Detection: An Examination of Wartime Russian Media (2022.findings-emnlp)

Copied to clipboard

Challenge: Information manipulation campaigns rely on textbased news and social media content, and NLP can be a valuable tool in combating them.
Approach: They propose to use a dataset to examine the use of NLP in public opinion manipulation campaigns in the 2022 Russia-Ukraine war.
Outcome: The proposed dataset contains 38M+ posts from Russian media outlets on Twitter and VKontakte, as well as public activity and responses, immediately preceding and during the 2022 Russia-Ukraine war.
An Empirical Methodology for Detecting and Prioritizing Needs during Crisis Events (2020.findings-emnlp)

Copied to clipboard

Challenge: Social media platforms such as Twitter contain a vast amount of information about the general public’s needs.
Approach: They propose to use Twitter to extract a list of needed resources and detecting sentences that specify who-needs-what resources.
Outcome: The proposed methods achieve 0.64 precision on a set of 1,000 annotated tweets and achieve 0.68 F1-score.
A Study in Contradiction: Data and Annotation for AIDA Focusing on Informational Conflict in Russia-Ukraine Relations (2022.lrec-1)

Copied to clipboard

Challenge: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives (AIDA) program . AIDA systems must extract entities, events, and relations from multimedia documents, aggregate that information across documents and languages, and produce multiple “hypotheses” about what has happened.
Approach: This paper describes data resources created for Phase 1 of the DARPA Active Interpretation of Disparate Alternatives program . the program aims to develop language technology that can help humans manage large volumes of conflicting information .
Outcome: The proposed corpus focuses on the domain of Russia-Ukraine relations and contains source data in English, Russian and Ukrainian . it is designed to support the development and evaluation of systems that extract entities, events, and relations from individual multimedia documents, aggregate the information across documents and languages, and produce multiple “hypotheses” about what has happened.
Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and Study (2024.lrec-main)

Copied to clipboard

Challenge: Mental illness can negatively impact individuals’ quality of life as it is considered one of the causes of years lived with disability and it is related to high suicide rates.
Approach: They collect first dataset of textual posts by same users before and after being diagnosed with depression and build multiple predictive models based on Transformers and BERT.
Outcome: The proposed model can be used to detect depression and suicidal thoughts in users who are not diagnosed with depression or suicide.
Emotion analysis and detection during COVID-19 (2022.lrec-1)

Copied to clipboard

Challenge: 3,000 English tweets labeled with emotions are used to predict emotions during crises . authors propose semi-supervised learning to bridge this gap .
Approach: They propose to use a dataset of 3,000 English tweets labeled with emotions . they propose semi-supervised learning to bridge this gap by analyzing unlabeled data .
Outcome: The proposed model can be used to predict emotions in the context of COVID-19 . the proposed model performs better than other models using unlabeled data .
EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian (2025.findings-emnlp)

Copied to clipboard

Challenge: **EmoBench-UA** is the first annotated dataset for emotion classification in Ukrainian texts.
Approach: They introduce **EmoBench-UA**, the first annotated dataset for emotion detection in Ukrainian texts.
Outcome: The first annotated dataset for emotion detection in Ukrainian texts is presented in this paper . the dataset was created through crowdsourcing using the Toloka.ai platform .
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI (2024.findings-emnlp)

Copied to clipboard

Challenge: Psychological trauma can manifest following various distressing events, but studies focus on a single aspect of trauma, often neglecting the transferability of findings across different scenarios.
Approach: They propose a language model that fine-tunes a single aspect of trauma to better predict traumatic events across domains.
Outcome: The proposed model outperforms large language models on trauma-related datasets . it also outperformed models on court data, counseling conversations, and forum posts .
Recognizing Social Cues in Crisis Situations (2024.lrec-main)

Copied to clipboard

Challenge: During natural disasters, observations of other people's behavior can play an essential role in a person's decision-making.
Approach: They propose a task to categorize social cues in tweets during crisis situations using an annotated dataset of 6,000 tweets.
Outcome: The proposed task is challenging for existing systems and a manual task is based on a dataset of 6,000 tweets labeled with eight social cue categories.
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show improvements on Reddit and Twitter data .
Approach: They propose to take advantage of Large Language Models (LLMs) to better identify user communities.
Outcome: The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations