Papers by Dustin Wright

12 papers
Generating Scientific Claims for Zero-Shot Scientific Fact Checking (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for scientific fact checking require domain expertise and time consuming.
Approach: They propose a new supervised method for generating claims from scientific sentences and a novel method for negating claims.
Outcome: The proposed method improves on existing methods on biomedical claims and negations.
Understanding Fine-grained Distortions in Reports of Scientific Findings (2024.findings-acl)

Copied to clipboard

Challenge: a fine-grained understanding of how scientific findings are reported is crucial, says a new study . a recent study found that tweets distort scientific findings more often than news reports .
Approach: They propose to annotate 1,600 scientific findings from academic papers paired with corresponding tweets . they also establish baselines for automatically detecting these characteristics .
Outcome: The proposed method outperforms few-shot prompting in detecting distortions in unpaired data.
Modeling Information Change in Science Communication with Semantically Matched Paraphrases (2022.emnlp-main)

Copied to clipboard

Challenge: Whether the media faithfully communicate scientific information has long been a core issue to the science community.
Approach: They propose to use the SCIENTIFIC PARAPHRASE AND INFORMATION CHANGE DATASET to identify paraphrased scientific findings annotated for degree of information change to enable large-scale tracking and analysis of information changes in science communication.
Outcome: The proposed dataset contains 6,000 scientific finding pairs extracted from news stories, social media discussions, and full texts of original papers.
Unstructured Evidence Attribution for Long Context Query Focused Summarization (2025.emnlp-main)

Copied to clipboard

Challenge: Existing systems struggle to copy and properly cite unstructured evidence, which also tends to be “lost-in-the-middle”.
Approach: They propose to extract unstructured evidence spans to improve the trustworthiness of large language models by citing unstructure . they propose to use this dataset as a training supervision for unstructure-based evidence summarization.
Outcome: The proposed pipeline generates more relevant and factually consistent evidence than baselines with no fine-tuning and fixed granularity evidence.
Transformer Based Multi-Source Domain Adaptation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve machine learning performance are mixed experts and domain adversarial training.
Approach: They investigate the problem of unsupervised multi-source domain adaptation . they combine predictions of multiple domain experts and combine them to induce a domain agnostic representation space .
Outcome: The proposed methods improve models' performance while limiting learning time.
Generating Label Cohesive and Well-Formed Adversarial Claims (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on adversarial triggers for fact checking models reveals weaknesses and flaws of models . universal adversarials often inadvertently invert the meaning of instances they are inserted in .
Approach: They propose a method for automatically generating highly potent, well-formed, label cohesive claims for FC using universal adversarial triggers.
Outcome: The proposed method maintains the directionality and semantic validity of the claim better than previous work on the FEVER dataset.
Stress Testing Factual Consistency Metrics for Long-Document Summarization (2026.acl-long)

Copied to clipboard

Challenge: Existing short-form summarization metrics struggle with input length limitations and long-range dependencies.
Approach: They propose to evaluate the reliability of six widely used reference-free factuality metrics in the long-document setting by applying seven factually-preserving perturbations to summaries.
Outcome: The proposed short-form summarization metrics struggle with long-range dependencies and input length limitations.
Semi-Supervised Exaggeration Detection of Health Science Press Releases (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that news media exaggerate scientific papers by exagging their findings.
Approach: They propose a method to detect when a news article has exaggerated a scientific finding . they use annotated press release/abstract pairs to compare machine learning models .
Outcome: The proposed method outperforms PET and supervised learning on a multi-task version of Pattern Exploiting Training.
CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Scientific document understanding is challenging due to the highly domain specific nature of scientific language.
Approach: They propose a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from extracted scientific documents.
Outcome: The proposed model improves on a paragraphlevel contextualized sentence labelling model based on Longformer . the model shows a 5 F1 point improvement over SciBERT which considers only individual sentences .
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
Claim Check-Worthiness Detection as Positive Unlabelled Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: a unified approach to claim check-worthiness detection is a critical component of fact checking systems.
Approach: They propose a unified approach which corrects for misinformation by positive unlabelled learning . they propose citation needed detection from Wikipedia and a ranking task which is a critical component of automatic fact checking systems.
Outcome: The proposed method outperforms the state of the art in two of the three tasks studied in English.
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings.
Approach: They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations.
Outcome: The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations