Papers by Aoife Cahill

8 papers
Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election (2025.coling-main)

Copied to clipboard

Challenge: Systematically organizing and geotagging large amounts of crowdsourced information requires substantial manual effort, often led by volunteers.
Approach: They present a dataset of 14k citizen reports related to the 2022 Kenyan General Election . they investigate whether language models can assist in scalably categorizing and geotagging reports .
Outcome: The proposed dataset aims to show whether language models can assist in categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space.
Don’t take “nswvtnvakgxpm” for an answer –The surprising vulnerability of automatic content scoring systems to adversarial input (2020.coling-main)

Copied to clipboard

Challenge: Automated content scoring systems can be used on short answer tasks to save human effort, but can invite cheating strategies such as writing irrelevant answers.
Approach: They generate adversarial answers for benchmark content scoring datasets based on different methods of increasing sophistication and examine countermeasures such as adversarials.
Outcome: The proposed methods show that even simple methods can reduce content scoring performance but do not solve the problem.
A New Task and Dataset on Detecting Attacks on Human Rights Defenders (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets for event extraction cannot extract information from textual sources.
Approach: They propose to use crowdsourced annotations on 500 online news articles to train and evaluate baseline models to predict annotated characteristics.
Outcome: The proposed dataset includes crowdsourced annotations on 500 online news articles and includes fine-grained information about the type and location of attacks, as well as information about victims.
HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid (2024.findings-emnlp)

Copied to clipboard

Challenge: Humanitarian organizations can analyze data to discover trends, gather aggregated insights, manage security risks, and inform advocacy and funding proposals.
Approach: They present a dataset comprising news articles in three languages containing instances of different types of violent incidents categorized by the humanitarian sector they impact.
Outcome: The proposed framework can be used to identify violent incidents and identify their impact on humanitarian operations.
Supporting Spanish Writers using Automated Feedback (2021.naacl-demos)

Copied to clipboard

Challenge: a tool that provides automated feedback to students studying Spanish provides useful results.
Approach: They propose a tool that provides automated feedback to students studying Spanish writing . the tool is made freely available via a free Google Docs add-on .
Outcome: The tool provides automated feedback to students studying Spanish writing . a small user study with 13 students in Mexico shows the tool is helpful .
Automated Scoring: Beyond Natural Language Processing (C18-1)

Copied to clipboard

Challenge: In this paper, we argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond competitive shared tasks.
Approach: They argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond standard competitive shared tasks . they argue that it is essential for us as NLP researchers to understand and incorporate these perspectives in our research and work towards a mutually satisfactory solution .
Outcome: The proposed approach is based on the findings of a recent conference on automated scoring.
Atypical Inputs in Educational Applications (N18-3)

Copied to clipboard

Challenge: atypical characteristics of some responses make it difficult for an automated scoring system to assign a valid score . a typical spoken response with a lot of background noise may suffer from frequent errors in automated speech recognition .
Approach: They propose a pipeline that detects and processes non-scorable responses at run-time . they also propose linguistic filtering models for spoken responses in language tests .
Outcome: The proposed pipeline detects and processes non-scorable responses at run-time and evaluates them for spoken responses in language proficiency assessment.
CEHA: A Dataset of Conflict Events in the Horn of Africa (2025.coling-main)

Copied to clipboard

Challenge: Existing datasets categorizing conflict events do not cover all of the fine-grained types of conflict relevant to areas like the Horn of Africa.
Approach: They propose to use online news articles to categorize violent conflict events . they propose to extract event-relevance and event-types from 500 English event descriptions .
Outcome: The proposed dataset categorizes conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations