Papers by Tomas Pfister

25 papers
SQLPrompt: In-Context Text-to-SQL with Minimal Labeled Data (2023.findings-emnlp)

Copied to clipboard

Challenge: Text-to-SQL aims to automate the process of generating SQL queries on a database from natural language text.
Approach: They propose a method to improve few-shot prompting capabilities of Text-to-SQL for Large Language Models (LLMs) they propose 'SQlPrompt' which aims to diversify the SQL proposals during consistency selection with different prompt designs and foundation models.
Outcome: The proposed method outperforms previous approaches for in-context learning with zero labeled data by a large margin, closing the gap with finetuning state-of-the-art with thousands of labeles.
Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation (2025.acl-long)

Copied to clipboard

Challenge: Large language models have been shown to be effective in multi-turn interactions . however, their performance may be limited in complex, multi-turned interactions involving users and multiple tools.
Approach: They propose a framework for synthesizing high-quality training trajectories to enhance the function calling capability of large language model agents in multi-turn conversations with humans.
Outcome: The proposed model outperforms the teacher model by 68.01 on BFCL-v3 and 73.30 on ToolQuery.
QueryForm: A Simple Zero-shot Form Entity Query Framework (2023.findings-acl)

Copied to clipboard

Challenge: Form-like document understanding is a key yet under-investigated problem . endlessly training specialized models on new document types is not scalable in many practical scenarios.
Approach: They propose to use large-scale query-entity pairs generated from form-like webpages to pre-train QueryForm.
Outcome: The proposed framework sets state-of-the-art average F1 score on XFUND and Payment benchmarks.
Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have enabled autonomous agents with complex reasoning and task-fulfillment capabilities using a wide range of tools.
Approach: They propose an unsupervised tool retrieval method that leverages LLM’s query understanding capabilities to extract key tool-related context and underlying intents from user queries.
Outcome: The proposed method significantly outperforms state-of-the-art tools in single-tool and multi-tool scenarios, all within a fully unsupervised setting.
TextGenSHAP: Scalable Post-Hoc Explanations in Text Generation with Long Documents (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are difficult to explain and understand due to long input contexts and autoregressive output generation.
Approach: They propose a post-hoc explanation method which incorporates LLM-specific techniques.
Outcome: The proposed method improves retrieval recall and prediction accuracy significantly on open-domain question answering benchmarks.
SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback (2026.findings-eacl)

Copied to clipboard

Challenge: High-quality, complex question-answer pairs are pivotal for training and evaluating capable deep search agents.
Approach: They propose a pipeline that generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level.
Outcome: The proposed pipeline generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level.
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation (2024.naacl-long)

Copied to clipboard

Challenge: Large language models generate "hallucinated" answers that are not factual . despite their widespread adoption, they can generate plausiblesounding but nonfactual information.
Approach: They propose a framework that tunes large language models to self-ground claims and provide citations to retrieved documents.
Outcome: The proposed framework generates superior grounded responses with more accurate citations compared to prompting-based approaches and post-hoc citing-based methods.
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction (2022.acl-long)

Copied to clipboard

Challenge: Form-like document understanding is a surging research topic due to its practical applications . form documents have unique challenges stemming from their structural characteristics .
Approach: They propose a structure-aware sequence model that leverages spatial relationships between tokens in a form for more precise attention score calculation.
Outcome: The proposed model outperforms existing methods with a more compact model size and less pre-training data.
Search-Adaptor: Embedding Customization for Information Retrieval (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to embed text in large language models are limited to zero-shot setups and can be integrated with any LLM.
Approach: They propose a method for customizing LLMs for information retrieval by modifying the embeddings generated by pre-trained LLM models and can be integrated with any LLM.
Outcome: The proposed method improves performance on English, multilingual, and multimodal retrieval datasets by 5% over 14 BEIR datasets.
Reverse Thinking Makes LLMs Stronger Reasoners (2025.naacl-long)

Copied to clipboard

Challenge: Reverse-Enhanced Thinking (RevThink) is a framework for large language models to perform reverse thinking.
Approach: They propose a framework for enhancing forward-backward reasoning by collecting data from a teacher model and employing three objectives to train a student model in a multi-task learning fashion.
Outcome: The proposed framework outperforms a fine-tuning method trained on 10x more forward reasoning on 12 datasets covering commonsense, math, and logical reasoning.
Found in the middle: Calibrating Positional Attention Bias Improves Long Context Utilization (2024.findings-acl)

Copied to clipboard

Challenge: Large language models struggle to capture relevant information located in the middle of their input.
Approach: They propose a calibration mechanism that allows the model to attend to contexts faithfully according to their relevance even when they are in the middle.
Outcome: The proposed calibration mechanism mitigates this positional bias and improves retrieval-augmented generation performance.
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to long-term dialogue memory management fail to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations.
Approach: They propose a mechanism that integrates forward- and backward-looking reflections into a personalized memory bank for effective future retrieval.
Outcome: The proposed mechanism outperforms state-of-the-art benchmarks on a long-term dialogue memory model.
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for natural planning lack constraint-guided iterative verification and adaptive selection . a recent study found that LLMs are not good at such planning.
Approach: They propose a model-agnostic and easily scalable agent framework with three key components: constraint, verification, and selection agents.
Outcome: The proposed framework improves inference-time algorithms on NATURAL PLAN and OlympiadBench benchmarks.
DocLens: A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to localizing evidence from long visual documents fail on a fundamental challenge: evidence localization.
Approach: They propose a tool-augmented multi-agent framework that “zooms in” on evidence like a lens.
Outcome: The proposed framework achieves state-of-the-art performance on MMLongBench-Doc and FinRAGBench-V, surpassing even human experts.
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes (2023.findings-acl)

Copied to clipboard

Challenge: Deploying large language models (LLMs) is difficult because they are memory inefficient and compute-intensive for practical applications.
Approach: They propose a mechanism that fine tunes or distills small models that outperform LLMs . they use human labels to fine tune models or LLM-generated labels to train models .
Outcome: The proposed method outperforms LLMs by using fewer training examples compared to few-shot prompted models using substantially smaller model sizes.
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to generate grounded responses are prone to errors due to the irrelevancy of input documents.
Approach: They propose a framework that leverages the insight that a robust grounded response should be consistent with information derived solely from its cited sources.
Outcome: Experiments on three open-domain question-answering datasets show that the proposed framework improves performance by 1.5% to 7% without any model fine-tuning.
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that decomposing complex problems into simple subtasks has significantly boosted the performance of large language models (LLMs).
Approach: They propose a unified post-training framework that distills synthetic task decompositions and fine-tunes smaller LLMs via supervised and reinforcement-learning objectives to improve complex reasoning.
Outcome: The proposed framework outperforms strong baselines on GSM8k and MATH benchmarks and shows that it can improve generalization capabilities on out-of-domain datasets.
Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive capabilities in many tasks, including natural language understanding and generation.
Approach: They propose a framework for adaptation with self-evaluation to improve selective prediction performance of large language models.
Outcome: The proposed framework outperforms state-of-the-art selective prediction methods on QA datasets and improves the AUACC from 91.23% to 92.63% and AUROC from 74.61% to 80.25%.
Universal Self-Adaptive Prompting (2023.emnlp-main)

Copied to clipboard

Challenge: a hallmark of modern large language models is their impressive general zero-shot and few-shot abilities . however, zero- shot performances are weaker due to the lack of guidance and the difficulty of applying existing automatic prompt design methods in general tasks.
Approach: They propose an automatic prompt design approach specifically tailored for zero-shot learning that categorizes a possible NLP task into one of three possible task types and then uses a selector to select the most suitable queries and zero- shot model-generated responses as pseudo-demonstrations.
Outcome: The proposed approach is able to generalize ICL to zero-shot learning tasks while also allowing for a more efficient and efficient prompt design.
When One LLM Drools, Multi-LLM Collaboration Rules (2026.acl-long)

Copied to clipboard

Challenge: a single general-purpose LLM is not enough to produce a reliable output, argues this paper . a multi-LLM collaboration approach addresses reliability, democratization, and pluralism .
Approach: They argue that a single general-purpose LLM is not enough to produce a reliable output . they organize existing multi-LLM collaboration methods into a hierarchy based on access and information exchange .
Outcome: The proposed method addresses reliability, democratization, and pluralism challenges a single LLM fails to produce a reliable output.
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions (2024.emnlp-main)

Copied to clipboard

Challenge: Embeddings from Large Language Models (LLMs) have emerged as critical components in information retrieval applications.
Approach: They propose a tuning framework for the customization of LLM embeddings.
Outcome: The proposed framework reduces embedding dimensions while maintaining comparable performance levels.
Better Zero-Shot Reasoning with Self-Adaptive Prompting (2023.findings-acl)

Copied to clipboard

Challenge: Modern large language models (LLMs) have demonstrated impressive capabilities at sophisticated tasks, often through step-by-step reasoning similar to humans.
Approach: They propose a new method that uses a set of examples from the LLM zero-shot outputs to improve performance.
Outcome: The proposed method improves performance up to 15% compared to baselines and matches or exceeds few-shot baselines at a range of reasoning tasks.
FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target designs, or additional pre-training data.
Approach: They propose a centralized multimodal graph contrastive learning strategy to unify self-supervised pre-training for all modalities in one loss.
Outcome: The proposed model achieves state-of-the-art performance on FUNSD, CORD, SROIE and Payment benchmarks with a more compact model size.
ROPE: Reading Order Equivariant Positional Encoding for Graph-based Document Information Extraction (2021.acl-short)

Copied to clipboard

Challenge: Graph Convolutional Networks (GCNs) have limited ability to capture reading orders of given word-level node representations in a graph.
Approach: They propose a new positional encoding technique to capture word-level nodes in a graph.
Outcome: The proposed method improves existing GCNs with an 8.4% F1 score on two datasets and a large-scale payment dataset.
CodecLM: Aligning Language Models with Tailored Synthetic Data (2024.findings-naacl)

Copied to clipboard

Challenge: Recent work on generating diverse instructions and applying LLM to increase instruction complexity neglects downstream use cases.
Approach: They propose a framework for generating high-quality synthetic data for LLM alignment with different downstream instruction distributions and LLMs.
Outcome: Experiments on four open-domain instruction using the proposed framework validate the effectiveness of CodecLM over the current state-of-the-art.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations