Papers by Elliott Ash

18 papers
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback (2024.acl-long)

Copied to clipboard

Challenge: a growing body of work on learning from human feedback to align various aspects of machine learning systems with human values and preferences is focusing on the setting of fairness in content moderation.
Approach: They propose to use human feedback to determine how two comments should be treated in content moderation to learn about human values and preferences.
Outcome: The proposed approach is promising, as human preferences can often not be A: Some ladies like smaller men. B: Some men like smaller guys. Figure 1 shows that the proposed approach performs better for demographic intersections than a single classifier that gives equal weight to each annotation.
The Empirical Variability of Narrative Perceptions of Social Media Texts (2024.emnlp-main)

Copied to clipboard

Challenge: Identifying stories in social media texts provides a lens through which we can study how individuals and communities process and communicate experiences.
Approach: They construct a taxonomy of crowd workers’ varied and nuanced perceptions of storytelling by open-coding their free-text rationales.
Outcome: The proposed model shows that crowd workers disagree on categorical labels, free-text storytelling rationales, authorial intent, and more.
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators (2024.acl-long)

Copied to clipboard

Challenge: generative AI is a counter-measure to misinformation, but factual claim detection suffers from inconsistency in definitions and high cost of manual annotation.
Approach: They propose a framework that assists in the annotation of factual claims with the help of large language models.
Outcome: The proposed framework can be used to annotate factual claims with the help of large language models and can work with or without expert supervision.
DIRAS: Efficient LLM Annotation of Document Relevance for Retrieval Augmented Generation (2025.naacl-long)

Copied to clipboard

Challenge: RAG systems leave out important relevant information (low recall) and excessively related but irrelevant information (high precision) authors propose a manual annotation-free schema that can be used for RAGs with limited performance.
Approach: They propose a manual annotation-free schema that annotates unseen queries with calibrated relevance scores.
Outcome: Evaluators show that DIRAS can achieve GPT-4-level performance on annotating and ranking unseen (query, document) pairs.
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? (2026.eacl-long)

Copied to clipboard

Challenge: Variation in human annotation (i.e., disagreements) is common in NLP, but it is unclear whether it is possible to model this variation in LLMs.
Approach: They evaluate the influence of different reasoning settings on LLM disagreement modeling . RLVR-style reasoning degrades performance in disagreement modeling, they find .
Outcome: The proposed reasoning settings improve LLM disagreement modeling, while RLVR-style reasoning degrades it.
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure (2025.emnlp-main)

Copied to clipboard

Challenge: Embedding-based similarity metrics can be influenced by content dimensions and spurious attributes like the text’s source or language.
Approach: They propose a debiasing algorithm that removes observed confounders from encoder representations and removes them from the encoder.
Outcome: The proposed method improves on out-of-distribution benchmarks and on benchmarks, but performance is not affected.
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification (2025.emnlp-demos)

Copied to clipboard

Challenge: Social scientists often need to develop codebooks that can be reliable but require significant human effort.
Approach: They propose a mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language models.
Outcome: The proposed framework integrates human expertise with automatic annotation guided by large language models.
Aligning Large Language Models with Diverse Political Viewpoints (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models such as ChatGPT exhibit striking political biases . a recent study shows that chatbots exhibit progressive, liberal, and proenvironmental biase .
Approach: They propose to align large language models with 100,000 comments from candidates running for national parliament in Switzerland.
Outcome: The proposed model generates more accurate political viewpoints from Swiss parties compared to commercial models such as ChatGPT.
Uncovering and Categorizing Social Biases in Text-to-SQL (2023.acl-long)

Copied to clipboard

Challenge: Existing Text-to-SQL models are trained on clean, neutral datasets, such as Spider and WikiSQl, but these models contain social bias at different rates.
Approach: They propose to use data to map natural language utterances to SQL queries.
Outcome: The proposed model can contain social bias at different rates in the downstream Text-to-SQL task.
The Law and NLP: Bridging Disciplinary Disconnects (2023.findings-emnlp)

Copied to clipboard

Challenge: Legal practitioners and scholars have been slow to adopt tools from natural language processing (NLP) the legal system is experiencing an access to justice crisis, which could be partially alleviated with NLP.
Approach: They argue that legal practitioners are slow to adopt natural language processing (NLP) they argue that there is a disconnect between legal needs and NLP research .
Outcome: The proposed tasks bridge disciplinary disconnects and highlight interesting areas for legal NLP research that remain underexplored.
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering (2024.acl-long)

Copied to clipboard

Challenge: Evidence-Based QA has proved insufficiently faithful with Large Language Models . a typical application of LLMs is in Evidence-based Question Answering (QA).
Approach: They propose a data generation pipeline with automated data quality filters to fine-tune LLMs for better source quality and answer attributability.
Outcome: The proposed model can synthesize high-quality training and testing data at scale.
Where Do People Tell Stories Online? Story Detection Across Online Communities (2024.acl-long)

Copied to clipboard

Challenge: Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text.
Approach: They propose a toolkit to detect stories in online communities using an annotated reddit dataset and a codebook adapted to social media context.
Outcome: The proposed toolkit includes an annotation-rich dataset of 502 Reddit posts and comments . it also includes a codebook adapted to the social media context and models to predict storytelling at document and span levels.
LePaRD: A Large-Scale Dataset of Judicial Citations to Precedent (2024.acl-long)

Copied to clipboard

Challenge: Legal passage retrieval is a practice-oriented task that seeks to predict relevant passages from precedential court decisions given the context of a legal argument.
Approach: They present a dataset which aims to facilitate work on legal passage retrieval . they extensively evaluate various approaches and find classification-based retrieval works best .
Outcome: The proposed dataset aims to facilitate work on legal passage retrieval . it shows that classification-based retrieval seems to work best .
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments (2026.acl-long)

Copied to clipboard

Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag
Challenge: Apertus is a fully open suite of large language models (LLMs) designed to address responsibility shortcomings in today’s open model ecosystem, namely data responsibility and global representation.
Approach: They propose to release a fully open suite of large language models (LLMs) that address data responsibility and global representation shortcomings in today’s open model ecosystem.
Outcome: The proposed model is pretrained on openly available data and suppresses verbatim recall of data while retaining task performance.
Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing verification approaches, such as Process Reward Models, are computationally expensive and limited to specific domains.
Approach: They propose a transformer-based probe that uses internal states of frozen LLMs to estimate credibility of reasoning steps during generation.
Outcome: The proposed probes match or exceed PRMs that are up to 810 larger.
Revisiting Automated Topic Model Evaluation with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Topic models are an unsupervised dimensionality reduction technique that help organize large text collections.
Approach: They propose to use large language models to evaluate document output and determine optimal number of topics.
Outcome: The proposed model performs better on coherence ratings of word sets than on intrustion detection.
MemSum: Extractive Summarization of Long Documents Using Multi-Step Episodic Markov Decision Processes (2022.acl-long)

Copied to clipboard

Challenge: MemSum is a reinforcement-learning-based extractive summarizer that considers the text content of the sentence, the global context of the rest of the document, and the extraction history of the sentences that have already been extracted.
Approach: They propose a reinforcement-learning-based extractive summarizer that iteratively selects sentences from a broad set of information that would intuitively be used by humans.
Outcome: The proposed extractive summarizer is enriched with information on the extraction history and local, global, and historical information.
Measuring scalar constructs in social science with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Valid scalar measurement of skalar constructs is a fundamental task in text analysis.
Approach: They evaluate four approaches to measuring scalar constructs using large language models . pairwise comparisons produced better measurements than prompting LLMs, they say . validation of skalar measurement enables wide range of substantive applications in social science research .
Outcome: The proposed methods improve on pairwise comparisons and finetuning . the proposed methods can be used in social science research .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations