Papers by Aoife Cahill
Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election (2025.coling-main)
Copied to clipboard
Roberto Mondini, Neema Kotonya, Robert L Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Systematically organizing and geotagging large amounts of crowdsourced information requires substantial manual effort, often led by volunteers. |
| Approach: | They present a dataset of 14k citizen reports related to the 2022 Kenyan General Election . they investigate whether language models can assist in scalably categorizing and geotagging reports . |
| Outcome: | The proposed dataset aims to show whether language models can assist in categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space. |
Don’t take “nswvtnvakgxpm” for an answer –The surprising vulnerability of automatic content scoring systems to adversarial input (2020.coling-main)
Copied to clipboard
| Challenge: | Automated content scoring systems can be used on short answer tasks to save human effort, but can invite cheating strategies such as writing irrelevant answers. |
| Approach: | They generate adversarial answers for benchmark content scoring datasets based on different methods of increasing sophistication and examine countermeasures such as adversarials. |
| Outcome: | The proposed methods show that even simple methods can reduce content scoring performance but do not solve the problem. |
A New Task and Dataset on Detecting Attacks on Human Rights Defenders (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for event extraction cannot extract information from textual sources. |
| Approach: | They propose to use crowdsourced annotations on 500 online news articles to train and evaluate baseline models to predict annotated characteristics. |
| Outcome: | The proposed dataset includes crowdsourced annotations on 500 online news articles and includes fine-grained information about the type and location of attacks, as well as information about victims. |
HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid (2024.findings-emnlp)
Copied to clipboard
Hemank Lamba, Anton Abilov, Ke Zhang, Elizabeth Olson, Henry Dambanemuya, João Bárcia, David Batista, Christina Wille, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Humanitarian organizations can analyze data to discover trends, gather aggregated insights, manage security risks, and inform advocacy and funding proposals. |
| Approach: | They present a dataset comprising news articles in three languages containing instances of different types of violent incidents categorized by the humanitarian sector they impact. |
| Outcome: | The proposed framework can be used to identify violent incidents and identify their impact on humanitarian operations. |
Supporting Spanish Writers using Automated Feedback (2021.naacl-demos)
Copied to clipboard
Aoife Cahill, James Bruno, James Ramey, Gilmar Ayala Meneses, Ian Blood, Florencia Tolentino, Tamar Lavee, Slava Andreyev
| Challenge: | a tool that provides automated feedback to students studying Spanish provides useful results. |
| Approach: | They propose a tool that provides automated feedback to students studying Spanish writing . the tool is made freely available via a free Google Docs add-on . |
| Outcome: | The tool provides automated feedback to students studying Spanish writing . a small user study with 13 students in Mexico shows the tool is helpful . |
Automated Scoring: Beyond Natural Language Processing (C18-1)
Copied to clipboard
| Challenge: | In this paper, we argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond competitive shared tasks. |
| Approach: | They argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond standard competitive shared tasks . they argue that it is essential for us as NLP researchers to understand and incorporate these perspectives in our research and work towards a mutually satisfactory solution . |
| Outcome: | The proposed approach is based on the findings of a recent conference on automated scoring. |
Atypical Inputs in Educational Applications (N18-3)
Copied to clipboard
| Challenge: | atypical characteristics of some responses make it difficult for an automated scoring system to assign a valid score . a typical spoken response with a lot of background noise may suffer from frequent errors in automated speech recognition . |
| Approach: | They propose a pipeline that detects and processes non-scorable responses at run-time . they also propose linguistic filtering models for spoken responses in language tests . |
| Outcome: | The proposed pipeline detects and processes non-scorable responses at run-time and evaluates them for spoken responses in language proficiency assessment. |
CEHA: A Dataset of Conflict Events in the Horn of Africa (2025.coling-main)
Copied to clipboard
Rui Bai, Di Lu, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Existing datasets categorizing conflict events do not cover all of the fine-grained types of conflict relevant to areas like the Horn of Africa. |
| Approach: | They propose to use online news articles to categorize violent conflict events . they propose to extract event-relevance and event-types from 500 English event descriptions . |
| Outcome: | The proposed dataset categorizes conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. |