Papers by Luca Soldaini
Knowledge Transfer from Answer Ranking to Answer Generation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that Question Answering (QA) based on Answer Sentence Selection (AS2) can be improved by generating an improved answer from the top-k ranked answer sentences. |
| Approach: | They propose to train a GenQA model by transferring knowledge from a trained AS2 model . they use top ranked candidate as the generation target and next k top rated candidates as context . |
| Outcome: | The proposed model outperforms existing models on public and industrial datasets. |
Pre-training Transformer Models with Sentence-Level Objectives for Answer Sentence Selection (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for answer sentence selection (AS2) are not yet available for AS2 . |
| Approach: | They propose to incorporate paragraph-level semantics within and across documents to improve transformers for AS2 . they propose to use a dataset to predict whether two sentences are extracted from the same paragraph . |
| Outcome: | The proposed model outperforms baseline models on public and industrial datasets on three public and one industrial dataset. |
Cross-Lingual Open-Domain Question Answering with Answer Sentence Generation (2022.aacl-main)
Copied to clipboard
| Challenge: | Open-Domain Generative Question Answering has achieved impressive performance in English . combining document-level retrieval with answer generation can generate complete sentences . |
| Approach: | They propose an open-domain approach that combines document retrieval with answer generation to generate complete sentences in English . they propose a cross-lingual generative model that exploits passages written in multiple languages . |
| Outcome: | The proposed model outperforms answer sentence selection baselines for all 5 languages and monolingual pipelines for three out of five languages. |
PaperMage: A Unified Toolkit for Processing, Representing, and Manipulating Visually-Rich Scientific Documents (2023.emnlp-demo)
Copied to clipboard
Kyle Lo, Zejiang Shen, Benjamin Newman, Joseph Chang, Russell Authur, Erin Bransom, Stefan Candra, Yoganand Chandrasekhar, Regan Huff, Bailey Kuehl, Amanpreet Singh, Chris Wilhelm, Angele Zamarron, Marti A. Hearst, Daniel Weld, Doug Downey, Luca Soldaini
| Challenge: | Existing tools for working with scientific documents are limited and documents are often in difficult-to-use PDF formats. |
| Approach: | They propose an open-source Python toolkit for analyzing and processing visually-rich scientific documents. |
| Outcome: | PaperMage provides turn-key recipes for common scientific document processing use-cases. |
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens (2025.acl-demo)
Copied to clipboard
Jiacheng Liu, Taylor Blanton, Yanai Elazar, Sewon Min, Yen-Sung Chen, Arnavi Chheda-Kothary, Huy Tran, Byron Bischoff, Eric Marsh, Michael Schmitz, Cassidy Trier, Aaron Sarnat, Jenna James, Jon Borchardt, Bailey Kuehl, Evie Yu-Yen Cheng, Karen Farley, Taira Anderson, David Albright, Carissa Schoenick, Luca Soldaini, Dirk Groeneveld, Rock Yuren Pang, Pang Wei Koh, Noah A. Smith, Sophie Lebrecht, Yejin Choi, Hannaneh Hajishirzi, Ali Farhadi, Jesse Dodge
| Challenge: | tracing language models' outputs back to training data is a problem because they are trained on text corpora with trillions of tokens . existing methods for tracers have not been scaled to work within this multi-trillion-token setting . |
| Approach: | They propose a system that traces language models' outputs verbatim back to training data . OLMOTRACE retrieves documents from the model's training data that contain exact matches . |
| Outcome: | The proposed system can find verbatim matches between LM output and training data . it can be used to explore fact checking, hallucination, and creativity of language models . |
Ensemble Transformer for Efficient and Accurate Ranking Tasks: an Application to Question Answering Systems (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large transformer models are expensive and slow to use in many applications. |
| Approach: | They propose an efficient neural network to distill large transformers into a single smaller model. |
| Outcome: | The proposed model outperforms existing models on English datasets . it outperformed existing models with 2.7 more parameters and 2.5 slower . |
MathFish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula (2024.findings-emnlp)
Copied to clipboard
| Challenge: | pedagogical experts spend months reviewing published math problems to ensure that they align with critical skills or concepts. |
| Approach: | They propose a novel approach for evaluating language models' mathematical abilities by combining a dataset of 385 fine-grained descriptions of K-12 math skills and concepts with 9.9K math problems labeled with these standards. |
| Outcome: | The proposed model can discern skills and concepts enabled by math content, and it can be used to assess language models' mathematical abilities. |
Answer Generation for Retrieval-based Question Answering Systems (2021.findings-acl)
Copied to clipboard
| Challenge: | Question Answering systems are a core component of many commercial applications . answer sentence selection (AS2) models are trained to select the best answer sentence . |
| Approach: | They propose to train a sequence to sequence transformer model to generate an answer from a set of candidates. |
| Outcome: | The proposed model improves accuracy by 32 points over the state-of-the-art model on English AS2 datasets. |
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students’ Hand-Drawn Math Images (2025.naacl-long)
Copied to clipboard
| Challenge: | DrawEduMath examines the ability of vision language models to handle real-world math problems, such as those encountered in classrooms and tutoring sessions. |
| Approach: | They present DrawEduMath, an English-language dataset of 2,030 images of students’ handwritten responses to math problems. |
| Outcome: | The proposed model can be used to evaluate teachers' QA pairs and 44,362 synthetic QAs derived from teachers' descriptions. |
SMHD: a Large-Scale Resource for Exploring Online Language Usage for Multiple Mental Health Conditions (C18-1)
Copied to clipboard
| Challenge: | Existing methods to label mental health conditions are based on high-precision diagnosis patterns and carefully selected control users. |
| Approach: | They propose to use high-precision diagnosis patterns to identify self-reported diagnoses of nine different mental health conditions and obtain high-quality labeled data without manual labelling. |
| Outcome: | The proposed dataset is two orders of magnitude larger than the largest published similar resource. |
The Cascade Transformer: an Application for Efficient Answer Sentence Selection (2020.acl-main)
Copied to clipboard
| Challenge: | Recent research shows that transformer-based neural networks can greatly advance the state of the art over many natural language processing tasks. |
| Approach: | They propose a technique to adapt transformer-based models into a cascade of rankers. |
| Outcome: | The proposed technique reduces computation by 37% with almost no impact on accuracy on two English question answering datasets. |
Embedding Recycling for Language Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies on embedding recycling have not adequately account for overhead costs. |
| Approach: | They propose to reuse contextualized embeddings from previous runs to speed training and inference of future ones. |
| Outcome: | The proposed technique speeds training and inference with no impact on accuracy. |
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions (2025.naacl-long)
Copied to clipboard
Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, Luca Soldaini
| Challenge: | Modern language models (LMs) are capable of following long and complex instructions that enable a large and diverse set of user requests. |
| Approach: | They propose a dataset that contains an instruction evaluation benchmark and a training set to help IR models learn to follow instructions. |
| Outcome: | The proposed model improves after fine-tuning on a training set and rigorous instruction evaluation benchmark. |
Open Domain Multi-document Summarization: A Comprehensive Study of Model Brittleness under Retrieval (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Multi-document summarization (MDS) assumes a set of topic-related documents is provided as input. |
| Approach: | They formalize the task and bootstrap it using existing datasets, retrievers and summarizers. |
| Outcome: | The proposed method reduces the sensitivity of summarizers to imperfect retrieval, but is highly sensitive to other errors. |
The olmOCR Project: Building Fully Open OCR using VLMs (2026.acl-demo)
Copied to clipboard
| Challenge: | olmOCR is a fully open OCR system developed through iterative public releases and community feedback. |
| Approach: | They propose an open OCR system that combines a 7B vision-language model trained in two stages: finetuning and reinforcement learning with visual unit tests. |
| Outcome: | The proposed system achieves state-of-the-art performance among open systems and proprietary APIs at a fraction of the cost. |
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets (2024.findings-eacl)
Copied to clipboard
| Challenge: | Using large language models (LMs) for query or document expansion can improve generalization in information retrieval. |
| Approach: | They conduct the first comprehensive analysis of large language models (LMs) for query or document expansion. |
| Outcome: | The proposed expansions improve retrieval performance for weaker models but harm stronger models. |
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models' (LLMs) abilities are drawn from their pretraining data. however, decisions around what data is retained or removed during this initial stage are under-scrutinized. |
| Approach: | They ground web text, a popular pretraining data source, to its social and geographic contexts. |
| Outcome: | The results show that some quality classifiers act like topical domain filters, and langID overlook English content from some regions of the world. |
A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific Documents (2023.emnlp-main)
Copied to clipboard
| Challenge: | snippets are not meant to be read outside their original document. |
| Approach: | They propose a framework that decomposes the task into three stages: question generation, question answering, and rewriting. |
| Outcome: | The proposed framework decomposes the task into three stages: question generation, question answering, and rewriting. |
Multi-task Learning of Spoken Language Understanding by Integrating N-Best Hypotheses with Hierarchical Attention (2020.coling-industry)
Copied to clipboard
| Challenge: | Existing methods to integrate hypotheses into speech recognition systems are noisy and can cause information loss. |
| Approach: | They propose to integrate hypotheses into multi-task learning and transfer learning to improve performance. |
| Outcome: | The proposed model improves domain and intent classification by 19% and 37% compared to current methods . the proposed model could recover transcription and rewrite the query for a better understanding . |
Modeling Context in Answer Sentence Selection Systems on a Latency Budget (2021.eacl-main)
Copied to clipboard
| Challenge: | Current AS2 models score question-answer pairs individually, ignoring any information from the document each potential answer was extracted from. |
| Approach: | They propose an approach to efficiently incorporate contextual information into AS2 models . they use unsupervised similarity techniques to extract relevant sentences from source document . |
| Outcome: | The proposed approach improves 6% to 11% over state-of-the-art in AS2 with minimal latency. |
KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used as conversational agents. |
| Approach: | They construct a dataset of knowledge-intensive writing instructions to evaluate LLMs' ability to follow user instructions. |
| Outcome: | The proposed model fails to integrate new information into an existing answer and perform precise and unambiguous edits. |
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature (2025.emnlp-main)
Copied to clipboard
David Wadden, Kejian Shi, Jacob Morrison, Alan Li, Aakanksha Naik, Shruti Singh, Nitzan Barzilay, Kyle Lo, Tom Hope, Luca Soldaini, Shannon Zejiang Shen, Doug Downey, Hannaneh Hajishirzi, Arman Cohan
| Challenge: | ScIRIFF is the only entirely expert-written instruction-following dataset for scientific literature understanding . it features complex instructions with long input contexts, detailed task descriptions, and structured outputs. |
| Approach: | They present a dataset of 137K instruction-following instances for training and evaluation . they finetuned large language models using a mix of general domain and ScIRIFF instructions . |
| Outcome: | The proposed dataset shows that on nine out-of-distribution held-out tasks, the model performs better than baselines trained on general domain instructions. |
Paragraph-based Transformer Pre-training for Multi-Sentence Inference (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained transformers perform poorly for multi-candidate inference tasks. |
| Approach: | They propose a pre-training objective that models paragraph-level semantics across multiple input sentences. |
| Outcome: | The proposed model outperforms existing models on three AS2 and one fact verification datasets. |