Papers by Joel Tetreault
Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election (2025.coling-main)
Copied to clipboard
Roberto Mondini, Neema Kotonya, Robert L Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Systematically organizing and geotagging large amounts of crowdsourced information requires substantial manual effort, often led by volunteers. |
| Approach: | They present a dataset of 14k citizen reports related to the 2022 Kenyan General Election . they investigate whether language models can assist in scalably categorizing and geotagging reports . |
| Outcome: | The proposed dataset aims to show whether language models can assist in categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space. |
XLTime: A Cross-Lingual Knowledge Transfer Framework for Temporal Expression Extraction (2022.findings-naacl)
Copied to clipboard
| Challenge: | Temporal Expression Extraction (TEE) is essential for understanding time in natural language. |
| Approach: | They propose a framework for multilingual Temporal Expression Extraction that leverages pre-trained language models to prompt cross-language knowledge transfer from English to non-English languages. |
| Outcome: | The proposed framework outperforms the existing SOTA methods on French, Spanish, Portuguese, and Basque by large margins. |
The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation (2020.coling-main)
Copied to clipboard
| Challenge: | appositives are phrases that appear next to a noun phrase and serve an explicative function. |
| Approach: | They propose a more realistic end-to-end definition of appositive generation with a dataset that spans four languages and two entity types. |
| Outcome: | The proposed model is non-trivial and leaves plenty of room for improvement. |
Event Extraction as Question Generation and Answering (2023.acl-short)
Copied to clipboard
| Challenge: | Recent work on Event Extraction addresses the error propagation issue found in token-based classification approaches. |
| Approach: | They propose a Question Generation (QG) model that generates questions that leverage contextual information instead of fixed templates. |
| Outcome: | The proposed model outperforms all previous single-task-based models on the ACE05 English dataset. |
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
Personalizing Grammatical Error Correction: Adaptation to Proficiency Level and L1 (D19-55)
Copied to clipboard
| Challenge: | Grammar error correction systems have become ubiquitous in a variety of software applications, but little is known about how to efficiently personalize them to the user’s characteristics, such as proficiency level and first language. |
| Approach: | They propose to adapt a general purpose neural GEC system to the proficiency level and the first language of a writer, using only a few thousand annotated sentences. |
| Outcome: | The proposed system improves on adapting to proficiency level and first language . the results are the broadest of its kind, covering five proficiency levels and twelve different languages. |
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics (2023.acl-long)
Copied to clipboard
Liang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes
| Challenge: | Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types. |
| Approach: | They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs. |
| Outcome: | The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors. |
An Exploration of Post-Editing Effectiveness in Text Summarization (2022.naacl-main)
Copied to clipboard
Vivian Lai, Alison Smith-Renner, Ke Zhang, Ruijia Cheng, Wenjuan Zhang, Joel Tetreault, Alejandro Jaimes-Larrarte
| Challenge: | Automated summarization methods are efficient but can suffer from low quality. |
| Approach: | They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience. |
| Outcome: | The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text. |
Rhetoric, Logic, and Dialectic: Advancing Theory-based Argument Quality Assessment in Natural Language Processing (2020.coling-main)
Copied to clipboard
| Challenge: | Existing work on argument quality (AQ) focuses on overall quality, but there is no large-scale theory-based corpus and corresponding computational models. |
| Approach: | They propose to use a large-scale English multi-domain argumentative writing corpus annotated with theory-based AQ scores to assess argument quality. |
| Outcome: | The proposed methods improve argument quality in three domains and can be used as strong baselines for future work. |
This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation (P19-1)
Copied to clipboard
| Challenge: | Existing research tracks on email use focus on email summarization, email keyword extraction and action detection. |
| Approach: | They propose to use email body to automatically generate an email subject line from the email body. |
| Outcome: | The proposed method outperforms baselines and state-of-the-art systems in the evaluation of human and automatic metrics. |
Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer (N18-1)
Copied to clipboard
| Challenge: | a lack of training and evaluation datasets, benchmarks and automatic metrics has blocked progress in this field. |
| Approach: | They propose to use a grammarly's Yahoo Answers Formality corpus to create the largest corpus for a particular style . they also propose to apply machine translation metrics to the task . |
| Outcome: | The proposed model can be used to train and evaluate a text in a particular style . the proposed model is based on the existing model and can be applied to other tasks . |
A New Task and Dataset on Detecting Attacks on Human Rights Defenders (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for event extraction cannot extract information from textual sources. |
| Approach: | They propose to use crowdsourced annotations on 500 online news articles to train and evaluate baseline models to predict annotated characteristics. |
| Outcome: | The proposed dataset includes crowdsourced annotations on 500 online news articles and includes fine-grained information about the type and location of attacks, as well as information about victims. |
Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer (2021.naacl-main)
Copied to clipboard
| Challenge: | XFORMAL benchmarks formal reformulations of informal text in Brazilian Portuguese, French, and Italian . most work on style transfer within English, while covering different languages has received disproportional interest. |
| Approach: | They create a benchmark of multiple formal reformulations of informal text in Brazil, Brazil, and Italy. |
| Outcome: | XFORMAL benchmarks formal reformulations of informal text in Brazilian Portuguese, French, and Italian . results show that state-of-the-art approaches perform close to simple baselines . |
Mapping the Design Space of Human-AI Interaction in Text Summarization (2022.naacl-main)
Copied to clipboard
| Challenge: | Automated text summarization systems involve humans for preparing data or evaluating model performance, yet, there is no systematic understanding of human-AI interactions and how to design for them. |
| Approach: | They conducted a systematic literature review of 70 papers and designed prototypes for each interaction. |
| Outcome: | The proposed design considerations were based on the results of a systematic literature review of 70 papers and interviews with 16 users. |
Defining a New NLP Playground (2023.findings-emnlp)
Copied to clipboard
Sha Li, Chi Han, Pengfei Yu, Carl Edwards, Manling Li, Xingyao Wang, Yi Fung, Charles Yu, Joel Tetreault, Eduard Hovy, Heng Ji
| Challenge: | Recent explosion of performance of large language models (LLMs) has changed the field more abruptly and seismically than any other shift in the field’s 80 year history. |
| Approach: | They propose 20+ PhD-dissertation-worthy research directions to define a new NLP playground by combining theoretical analysis, new and challenging problems, learning paradigms and interdisciplinary applications. |
| Outcome: | The proposed research will cover theoretical analysis, new and challenging problems, learning paradigms and interdisciplinary applications. |
HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid (2024.findings-emnlp)
Copied to clipboard
Hemank Lamba, Anton Abilov, Ke Zhang, Elizabeth Olson, Henry Dambanemuya, João Bárcia, David Batista, Christina Wille, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Humanitarian organizations can analyze data to discover trends, gather aggregated insights, manage security risks, and inform advocacy and funding proposals. |
| Approach: | They present a dataset comprising news articles in three languages containing instances of different types of violent incidents categorized by the humanitarian sector they impact. |
| Outcome: | The proposed framework can be used to identify violent incidents and identify their impact on humanitarian operations. |
Journalistic Guidelines Aware News Image Captioning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that JoGANIC outperforms state-of-the-art methods for image caption generation. |
| Approach: | They propose a method to generate descriptive and informative captions for news article images . they leverage the structure of captions to improve the generation quality and guide their representation . |
| Outcome: | The proposed method outperforms state-of-the-art methods on two large-scale datasets. |
CrisisLTLSum: A Benchmark for Local Crisis Event Timeline Extraction and Summarization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Timeline extraction and abstractive summarization are critical tasks for leveraging large numbers of social media posts about events. |
| Approach: | They propose to build a semi-automated cluster-then-refine algorithm to extract local crisis event timelines from Twitter. |
| Outcome: | The proposed approach performs better than human models on extraction and summarization tasks. |
Multi-View Source Ablation for Faithful Summarization (2023.findings-eacl)
Copied to clipboard
| Challenge: | MuFaSSa is a metric for evaluating faithfulness of abstractive summaries . it uses different strategies to remove information from source document to form multiple ablated views . |
| Approach: | They propose a metric for evaluating faithfulness of abstractive summaries using multiple ablated views. |
| Outcome: | The proposed metric outperforms existing models on summarization tasks and human-annotated faithfulness labels. |
Dialogue Act Classification with Context-Aware Self-Attention (N19-1)
Copied to clipboard
| Challenge: | Recent work in Dialogue Act classification has treated the task as a sequence labeling problem using hierarchical deep neural networks. |
| Approach: | They propose a hierarchical deep neural network to model different levels of utterance and dialogue act semantics and use contextual dependencies to improve performance. |
| Outcome: | The proposed model improves on the Switchboard Dialogue Act Corpus while maintaining high accuracy. |
CEHA: A Dataset of Conflict Events in the Horn of Africa (2025.coling-main)
Copied to clipboard
Rui Bai, Di Lu, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel Tetreault, Alejandro Jaimes
| Challenge: | Existing datasets categorizing conflict events do not cover all of the fine-grained types of conflict relevant to areas like the Horn of Africa. |
| Approach: | They propose to use online news articles to categorize violent conflict events . they propose to extract event-relevance and event-types from 500 English event descriptions . |
| Outcome: | The proposed dataset categorizes conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. |
Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have shattered the ceiling of human-like text generation. |
| Approach: | They compared human-AI interaction types in LLM-assisted news headline generation to determine whether humans can best leverage them for writing. |
| Outcome: | The guiding and selecting model outputs added the most benefit with the lowest cost (in time and effort) Furthermore, AI assistance did not harm participants’ perception of control compared to freeform editing. |