HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid (2024.findings-emnlp)
Copied to clipboard
Hemank Lamba, Anton Abilov, Ke Zhang, Elizabeth Olson, Henry Dambanemuya, João Bárcia, David Batista, Christina Wille, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Humanitarian organizations can analyze data to discover trends, gather aggregated insights, manage security risks, and inform advocacy and funding proposals. |
| Approach: | They present a dataset comprising news articles in three languages containing instances of different types of violent incidents categorized by the humanitarian sector they impact. |
| Outcome: | The proposed framework can be used to identify violent incidents and identify their impact on humanitarian operations. |
Similar Papers
HumSet: Dataset of Multilingual Information Extraction and Classification for Humanitarian Crises Response (2022.findings-emnlp)
Copied to clipboard
Selim Fekih, Nicolo’ Tamagnone, Benjamin Minixhofer, Ranjan Shrestha, Ximena Contla, Ewan Oglethorpe, Navid Rekabsaz
| Challenge: | During humanitarian crises, a quick and accurate analysis of relevant data is critical to a timely and effective response. |
| Approach: | They introduce and release a multilingual dataset of humanitarian response documents annotated by experts in the humanitarian response domain. |
| Outcome: | The proposed dataset provides documents in three languages and covers a variety of humanitarian crises from 2018 to 2021 across the globe. |
MultiHumES: Multilingual Humanitarian Dataset for Extractive Summarization (2021.eacl-main)
Copied to clipboard
| Challenge: | a new multilingual summarization model is being developed to help humanitarian experts process large amounts of secondary data to derive situational awareness and guide decision-making. |
| Approach: | They propose to use multilingual documents and annotated snippets to improve extraction of secondary data for humanitarian response experts. |
| Outcome: | The proposed model provides multilingual documents with informative snippets that have been annotated by humanitarian analysts over the past four years. |
A New Task and Dataset on Detecting Attacks on Human Rights Defenders (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for event extraction cannot extract information from textual sources. |
| Approach: | They propose to use crowdsourced annotations on 500 online news articles to train and evaluate baseline models to predict annotated characteristics. |
| Outcome: | The proposed dataset includes crowdsourced annotations on 500 online news articles and includes fine-grained information about the type and location of attacks, as well as information about victims. |
A Dataset for Multi-lingual Epidemiological Event Extraction (2020.lrec-1)
Copied to clipboard
| Challenge: | Using the Web, we propose a corpus for information extraction and text classification. |
| Approach: | They propose to use a corpus for information extraction and natural language processing (NLP) tasks such as text classification. |
| Outcome: | The proposed corpus can be used for information extraction and natural language processing tasks such as text classification. |
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)
Copied to clipboard
Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Röttger
| Challenge: | Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect . |
| Approach: | They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks . |
| Outcome: | The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis . |
SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification (2021.findings-acl)
Copied to clipboard
| Challenge: | toxicity, hate speech, cyberbullying, and cyber-aggression are common themes in social media . authors present a dataset that is limited in size and biased towards offensive language . |
| Approach: | They present an expanded dataset that uses a taxonomy for offensive language identification . they show that using SOLID and OLID yields sizable performance gains . |
| Outcome: | The proposed dataset shows that it performs better than the OLID dataset for two different models. |
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge. |
| Approach: | They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain. |
| Outcome: | The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain. |
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)
Copied to clipboard
| Challenge: | Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders. |
| Approach: | They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content. |
| Outcome: | The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical. |
CEHA: A Dataset of Conflict Events in the Horn of Africa (2025.coling-main)
Copied to clipboard
Rui Bai, Di Lu, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Existing datasets categorizing conflict events do not cover all of the fine-grained types of conflict relevant to areas like the Horn of Africa. |
| Approach: | They propose to use online news articles to categorize violent conflict events . they propose to extract event-relevance and event-types from 500 English event descriptions . |
| Outcome: | The proposed dataset categorizes conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. |
Humanitarian Corpora for English, French and Spanish (2024.lrec-main)
Copied to clipboard
| Challenge: | et al., a leading database of humanitarian documents, compiled with ReliefWeb reports . documents selected with language identification and noise reduction techniques . authors present corpora of English, French and Spanish humanitarian documents . |
| Approach: | They present three corpora of English, French and Spanish humanitarian documents compiled with ReliefWeb reports . documents were tokenized, lemmatized, tagged by part of speech, and enriched with metadata . authors propose a project to develop a humanitarian dictionary with a focus on conceptual variation . |
| Outcome: | The corpora were compiled to satisfy the research needs of the Humanitarian Encyclopedia project with a focus on conceptual variation. |