Papers by Daphne Ippolito

16 papers
Automatic Detection of Generated Text is Easiest when Humans are Fooled (2020.acl-main)

Copied to clipboard

Challenge: Recent advances in neural language modelling make it possible to rapidly generate vast amounts of human-sounding text.
Approach: They compare decoding methods with popular sampling-based decoding strategies . they show that multi-sentence excerpts can fool expert human raters over 30% of the time .
Outcome: The proposed methods improve with longer excerpt length, but multi-sentence excerpts fool human raters over 30% of the time.
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
CAVA: A Tool for Cultural Alignment Visualization & Analysis (2024.emnlp-demo)

Copied to clipboard

Challenge: Using CAVA, researchers can analyze country-specific biases encoded in large language models.
Approach: They propose a visualization tool that allows users to identify biases in language models by adding country-based questions and models.
Outcome: The proposed tool can be used to analyze the cultural competencies of large language models across the dimension of geographic locales.
Chasing Random: Instruction Selection Strategies Fail to Generalize (2025.findings-naacl)

Copied to clipboard

Challenge: Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions.
Approach: They analyze popular selection strategies across different datasets and benchmarks to find out whether they generalize poorly.
Outcome: The proposed methods outperform random baselines and cost-performance trade-offs on the full dataset and a random subset.
ChatEval: A Tool for Chatbot Evaluation (N19-4)

Copied to clipboard

Challenge: open-domain dialog systems are difficult to evaluate due to lack of standardization and standardization in evaluation procedures.
Approach: They propose a framework for human evaluation of chatbots that augments existing tools . researchers can submit their trained models to the ChatEval web interface . reproducibility and model assessment for opendomain dialog systems is challenging .
Outcome: The proposed framework provides a web-based hub for researchers to compare their models with baselines and prior work.
The Case for a Single Model that can Both Generate Continuations and Fill-in-the-Blank (2022.findings-naacl)

Copied to clipboard

Challenge: a natural language generation system can be used to create text at the end of a passage . fill in the blank (FITB) is a task of inserting text into a specified position in a text .
Approach: They evaluate the feasibility of using a single model to perform both tasks . they show that models pre-trained with a FitB-style objective are capable of both tasks.
Outcome: The proposed model can perform both fill in the blank and continuation tasks.
Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence (2022.emnlp-main)

Copied to clipboard

Challenge: researchers have posited Dungeons and Dragons as a challenge problem to test systems on various language-related capabilities.
Approach: They frame Dungeons and Dragons specifically as a dialogue system challenge . they train a large language model to generate the next game turn, conditioning it on different information.
Outcome: The proposed game generates the next conversational turn and predicts the state of the game given the dialogue history.
Comparison of Diverse Decoding Methods from Conditional Language Models (P19-1)

Copied to clipboard

Challenge: Conditional language models can generate a diverse set of outputs, but for open-ended tasks, beam search is ill-suited to generating a set of diverse sequences.
Approach: They propose a method where we over-sample candidates and use clustering to remove similar sequences to achieve high diversity without sacrificing quality.
Outcome: The proposed method over-samples candidates and removes similar sequences to achieve high diversity without sacrificing quality.
CIE: Controlling Language Model Text Generations Using Continuous Signals (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to control language models with intent are brittle and hard to scale.
Approach: They propose to use a set of LMs to fine-tune to expect a control vector that is interpolated between a "low" and a 'high' token embedding.
Outcome: The proposed method can be finetuned to expect a control vector that is interpolated between a “low” and a ‘high” token embedding.
A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity (2024.naacl-long)

Copied to clipboard

Challenge: a large number of pretraining data design practices are under-documented, authors say . authors: strong performance of modern language models depends on selfsupervised pretraining .
Approach: They propose to pretrain models on data curated at different collection times . they find temporal shift between evaluation data and pretraining data leads to performance degradation .
Outcome: The results validate, quantify, and expose many undocumented intuitions about text pretraining . authors say this practice has outperformed other models in the field .
Toward Better Storylines with Sentence-Level Language Models (2020.acl-main)

Copied to clipboard

Challenge: Rather than modeling fluency, the sentence-level language model can focus on longer range dependencies, which are crucial for multi-sentence coherence.
Approach: They propose a sentence-level language model which selects the next sentence in a story from a finite set of fluent alternatives.
Outcome: The proposed model can focus on longer range dependencies, crucial for multi-sentence coherence.
RoFT: A Tool for Evaluating Human Detection of Machine-Generated Text (2020.emnlp-demos)

Copied to clipboard

Challenge: Existing studies on how humans perceive machine-generated text are limited due to the prohibitive cost of running human evaluation studies.
Approach: They propose a task to detect the boundary at which a text passage starts off human-written transitions to being machine-generated.
Outcome: The proposed system evaluates machine-generated news articles on a wide range of domains.
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting machine-generated text are often insufficiently robust and lack benchmark datasets.
Approach: They evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors using RAID benchmark datasets.
Outcome: The proposed detectors are fooled by adversarial attacks, repetition penalties, and unseen generative models.
A Recipe for Arbitrary Text Style Transfer with Large Language Models (2022.acl-short)

Copied to clipboard

Challenge: augmented zero-shot learning is a prompting method that allows large language models to perform zero-shoot text style transfer to arbitrary styles, without any model fine-tuning or exemplars in the target style.
Approach: They propose a prompting method that frames style transfer as a sentence rewriting task and requires only a natural language instruction.
Outcome: The proposed method is based on a large language model and is shown to perform on standard style transfer tasks and arbitrary transformations.
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Reliable multilingual evaluation is difficult and culturally appropriate evaluation is even harder to achieve.
Approach: They propose a multilingual evaluation framework that aims to mitigate these biases by improving translations and annotation practices.
Outcome: The proposed framework improves translation quality and cultural coverage and is culturally sensitive and culturally agnostic.
Deduplicating Training Data Makes Language Models Better (2022.acl-long)

Copied to clipboard

Challenge: Existing language modeling datasets contain near-duplicate examples and long repetitive substrings.
Approach: They develop tools that allow us to deduplicate existing language modeling datasets . they found that over 1% of the unprompted output of language models is copied verbatim .
Outcome: The proposed tools reduce train-test overlap, which affects over 4% of validation sets, and improve model accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations