Papers by Qingfu Zhu

31 papers
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query (2025.emnlp-main)

Copied to clipboard

Challenge: Existing KV cache eviction methods prune tokens using prefilling-stage attention scores, causing inconsistency with actual inference queries.
Approach: They propose a lookahead q-cache framework that generates low-cost pseudo lookaheaded queries to better approximate the true decoding-stage queries.
Outcome: The proposed framework outperforms existing methods on LongBench and Needle-in-a-Haystack benchmarks and can be flexibly combined to yield further improvements.
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format (2026.findings-acl)

Copied to clipboard

Challenge: Prior work showed that multiple reasoning formats outperform a single format when generating multiple answers.
Approach: They propose a method to measure reasoning error when generating multiple answers . they propose 'formatadapter' which generates and selects suitable reasoning formats .
Outcome: The proposed method achieves a 4.3% performance improvement over previous works on math and commonsense reasoning tasks.
MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQL (2025.coling-main)

Copied to clipboard

Challenge: Existing multi-hop retrieval of open-domain text-to-SQL tasks is not applicable due to the tendency to retrieve tables similar to those already retrieved but irrelevant to the question.
Approach: They propose a multi-hop table retrieval with removal task to retrieve unretrieved tables from open-domain text-to-SQL databases.
Outcome: The proposed method improves performance 5.7% over the previous state-of-the-art methods on open-domain text-to-SQL datasets.
Python is Not Always the Best Choice: Embracing Multilingual Program of Thoughts (2024.emnlp-main)

Copied to clipboard

Challenge: Program of Thoughts (PoT) is an approach characterized by its executable intermediate steps, which ensure the accuracy of the logical calculations in the reasoning process.
Approach: They propose a task and model agnostic approach which harnesses strength and diversity from various languages to achieve better performance across all tasks.
Outcome: The proposed approach outperforms Python Self-Consistency in almost all tasks and models and achieves comparable or superior performance on ChatGPT.
DAC: Decomposed Automation Correction for Text-to-SQL (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve text-to-SQL performance are hard to detect errors in SQL directly.
Approach: They propose to use decomposed correction to improve text-to-SQL performance . they first detect errors based on decompose subtasks, then use it to correct them .
Outcome: The proposed method improves text-to-SQL performance by 1.4% compared with previous methods .
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training (2024.emnlp-main)

Copied to clipboard

Challenge: Existing speculative decoding methods require additional model structure and training processes to assist the model for draft token generation.
Approach: They propose a make some noise training framework that introduces some noise at the input for the model to learn the denoising task.
Outcome: The proposed model improves inference speed by 2.3-2.7x times without compromising model performance.
Abacus-SQL: A Text-to-SQL System Empowering Cross-Domain and Open-Domain Database Retrieval (2025.acl-demo)

Copied to clipboard

Challenge: Existing text-to-SQL systems often lack retrieval capabilities for open-domain databases, requiring users to manually filter relevant databases.
Approach: They propose to use database retrieval technology to locate the required databases in an open-domain database environment and enhance system cross-domain transferability through data augmentation methods.
Outcome: The proposed system performs excellently in multi-turn text-to-SQL tasks, validating the proposed approach’s effectiveness.
Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes (2024.acl-long)

Copied to clipboard

Challenge: Numerical reasoning is an essential ability for NLP systems to handle numeric information.
Approach: They propose a numerical reasoning method that generates reliable reasoning processes by decomposing the answer formula and aim to train models to generate the process with synthesized data.
Outcome: The proposed method improves on all five datasets with an average improvement of 1.8% compared with baselines and gpt-3.5-turbo.
Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring (2025.naacl-long)

Copied to clipboard

Challenge: Existing black-box jailbreak methods often rely on model feedback . existing methods may be intercepted by content moderators during the search process .
Approach: They propose a method that guides malicious prompt construction by local training a mirror model of the target black-box model through benign data distillation.
Outcome: The proposed method achieves a 92% attack success rate and 80% stealth rate on a subset of AdvBench.
Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated exceptional performance in reasoning tasks, particularly in mathematics.
Approach: They propose a dynamic self-consistency framework that integrates System 1 and System 2 reasoning to improve token efficiency and latency.
Outcome: The proposed method outperforms existing methods, achieving up to 47% reduction in token consumption and 43% reduction in inference latency without significant performance loss.
Retrieval-Enhanced Adversarial Training for Neural Response Generation (P19-1)

Copied to clipboard

Challenge: Existing approaches to dialogue systems are labor-intensive and difficult to scale up.
Approach: They propose a Retrieval-Enhanced Adversarial Training method for neural response generation that leverages an adversarial training paradigm while taking advantage of N-best response candidates from a retrieval-based system to construct the discriminator.
Outcome: The proposed method outperforms the vanilla Seq2Seq model and conventional adversarial training approach on a large scale dataset.
Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards (2026.acl-long)

Copied to clipboard

Challenge: Existing efforts to generate static visualizations focus on static charts and interactive dashboards.
Approach: They propose a dashboard2code task that requires a model to explore an interactive dashboard, acquire feedback from its own interactions and generate code that reproduces the target dashboard.
Outcome: The proposed task is based on 180 carefully designed and manually verified dashboard–code pairs spanning three difficulty levels and covering eight common real-world interaction patterns.
Concise and Precise Context Compression for Tool-Using Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods suffer from key information loss and difficulty in adjusting the length of compressed sequences based on documentation lengths.
Approach: They propose two strategies for compressing tool documentation into concise and precise summary sequences for tool-using language models.
Outcome: The proposed approach achieves comparable performance to the upper-bound baseline under 16x compression ratio.
Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits (2025.coling-main)

Copied to clipboard

Challenge: Existing data on MBTI personality detection are based on self-reported labels and fail to capture the full range of population personality traits.
Approach: They construct a manually annotated MBTI personality detection dataset with soft labels under the guidance of psychologists and use them to identify the task.
Outcome: The MBTIBench is the first manually annotated MBti personality detection dataset with soft labels under the guidance of psychologists.
Improving Grammatical Error Correction via Contextual Data Augmentation (2024.findings-acl)

Copied to clipboard

Challenge: Increasing use of synthetic data due to inconsistent error distribution and noisy labels is limiting the use of these data.
Approach: They propose a method for augmentation of synthetic data with a more consistent error distribution.
Outcome: The proposed method outperforms strong baselines and achieves state-of-the-art with only a few synthetic data.
A Survey on Natural Language Processing for Programming (2024.lrec-main)

Copied to clipboard

Challenge: Natural language processing for programming is a field of NLP and software engineering . it is used to assist programming, and is increasingly prevalent for its effectiveness in improving productivity.
Approach: They propose to use NLP techniques to assist programming by obtaining a structure-based representation and a functionality-oriented algorithm.
Outcome: The proposed approach could relieve developers from laborious work while improving efficiency for non-professional users.
Self-Constructed Context Decompilation with Fined-grained Alignment Enhancement (2024.findings-emnlp)

Copied to clipboard

Challenge: Decompilation is the process of converting compiled code back into a high-level programming language for analysis when source code is unavailable.
Approach: They propose two methods to improve decompilation performance without fine-tuning and fine-grained alignment enhancement to achieve further improvements.
Outcome: The proposed methods achieved a Re-Executability performance improvement of approximately 3.90% on the Decompile-Eval benchmark, establishing a new state-of-the-art performance of 52.41%.
When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models (2026.acl-long)

Copied to clipboard

Challenge: Vision-Language-Action models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood.
Approach: They propose a step-wise inference-time intervention that aligns representations according to step language sensitivity, significantly improving performance under linguistic variation.
Outcome: The proposed model significantly improves performance under linguistic variation under non-English instructions under language-agnostic steps.
Scaling Laws for Code: A More Data-Hungry Regime (2026.acl-long)

Copied to clipboard

Challenge: Code Large Language Models (LLMs) are revolutionizing software engineering, but scaling laws are primarily analyzed on Natural Language (NL).
Approach: They fit Chinchilla law and Farsser law to test scaling laws for code . they find code is more data-hungry and requires higher data-to-parameter ratio .
Outcome: The proposed scaling laws show that the more expressive Farsser law offers greater accuracy and scales with model size.
Context-Sensitive Generation of Open-Domain Conversational Responses (C18-1)

Copied to clipboard

Challenge: Existing studies on single-turn conversation generation focus on coherence and context-sensitive generation of open-domain conversational responses.
Approach: They propose static and dynamic attention based approaches for context-sensitive generation of open-domain conversational responses.
Outcome: The proposed model outperforms all baselines on automatic and human evaluation on two public datasets.
RoT: Enhancing Table Reasoning with Iterative Row-Wise Traversals (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in reasoning large language models (RLLMs) have significantly enhanced reasoning capabilities, leading to brilliant performance on table reasoning.
Approach: They propose a method which performs iterative row-wise table traversal, allowing for reasoning extension and reflection-based refinement at each traversal.
Outcome: Experiments show that the proposed method outperforms RLLMs on WikiTableQuestions and TableBench by 4.3% and achieves state-of-the-art results with comparable models.
Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing Chart2code-related training datasets suffer from limited scale, limited type coverage, and inadequate complexity.
Approach: They propose to synthesize chart2code-related training datasets using web plotting code and chart images to address these challenges.
Outcome: The proposed dataset exhibits the greatest diversity and higher complexity compared to other open-source Chart2code related datasets.
SCITAT: A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types (2025.findings-acl)

Copied to clipboard

Challenge: Existing scientific question answering datasets lack diverse reasoning types and neglect relevance between tables and text.
Approach: They propose a scientific question answering benchmark for scientific tables and text with diverse reasoning types (SCITAT) to address these challenges, they propose QA benchmark which incorporates tables and texts to ensure that the questions encompass both tables and textes.
Outcome: The proposed benchmark improves by 4.1% over baselines on SCITAT.
Counterfactual Off-Policy Training for Neural Dialogue Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for open-domain dialogue generation suffer from data insufficiency . a potential response inferred in hindsight is called a counterfactual reasoning .
Approach: They propose to explore potential responses by counterfactual reasoning . given an observed response, the model automatically infers the outcome of an alternative policy that could have been taken .
Outcome: The proposed model outperforms the HRED model and conventional learning frameworks on the DailyDialog dataset.
Neural Stylistic Response Generation with Disentangled Latent Variables (2021.acl-long)

Copied to clipboard

Challenge: Existing parallel datasets for creating stylistic responses are not stylistically consistent.
Approach: They propose to disentangle the content and style in latent space by diluting sentence-level information in style representations.
Outcome: The proposed approach achieves a higher BERT-based style intensity score and comparable BLEU scores, compared with baselines.
OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence.
Approach: They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included .
Outcome: The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model.
Exploring Hybrid Question Answering via Program-based Prompting (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to question answering over heterogeneous data are limited due to large scale of information and organic coupling of heterogenous data.
Approach: They propose a program-based prompting framework for hybrid question answering tasks . it integrates various functions to perform hybrid information-seeking over data .
Outcome: The proposed framework surpasses baseline systems and achieves the best performance under the fewshot settings.
Tag-Evol: Achieving Efficient Instruction Evolving via Tag Injection (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods rely on a fixed set of strategies to evolve, which requires manual design and is monolithic in form.
Approach: They propose a method that uses diverse and specific knowledge tags to achieve controlled evolution by injecting different combinations of tags into original instructions.
Outcome: The proposed method generates better evolved data than existing methods and is more diverse and challenging.
Improving Demonstration Diversity by Human-Free Fusing for Text-to-SQL (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have explored selecting relevant demonstrations from a human-labeled demonstration pool, but these methods lack diversity and incur high labeling costs.
Approach: They propose a method that iteratively fuses demonstrations to create a diverse demonstration pool based on human labeling or even from scratch with LLMs, reducing labeling costs.
Outcome: The proposed method achieves an average improvement of 2.1% based on existing labeling and 5.5% from scratch on mainstream datasets.
Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate only one token at each decoding step, leading to high latency.
Approach: They propose a speculative decoding paradigm that stores tokens in an adjacency matrix and employs a breadth-first-search algorithm to construct a draft tree.
Outcome: The proposed method outperforms existing train-free methods by 30% and even a training method by 25%.
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing TATQA datasets are limited to English, leading to drawbacks . existing datasets overlook challenges of multilingual TAT-QA and do not reflect real-world multilingual scenarios .
Approach: They propose a multilingual TATQA dataset that can be translated into 10 languages . they use data from 3 mainstream TATQ datasets and analyze the results .
Outcome: The proposed dataset outperforms other baselines by an average of 3.3 .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations