Papers by Lefei Zhang

14 papers
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for decoding large language models generate one token per step, causing high inference latency.
Approach: They propose a method that integrates retrieved exact patterns with logit-driven future cues.
Outcome: Experiments on Spec-Bench, HumanEval, and MGSM-ZH show that RACER outperforms training-free methods and accelerates inference.
Intention Analysis Makes LLMs A Good Jailbreak Defender (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to align large language models with human values overlook the intrinsic nature of jailbreaks, which limits their effectiveness in complex scenarios.
Approach: They propose a simple yet highly effective defense strategy, i.e., Intention Analysis (IA). They show that IA suppresses LLM’s tendency to follow jailbreak prompts, thereby enhancing safety.
Outcome: The proposed strategy reduces harmfulness of LLMs and outperforms GPT-3.5 in attack success rate.
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show ToM outperforms existing divide-and-conquer frameworks . RAG relies on similarity-based rankings to retrieve and reason over chunks based on logical coherence .
Approach: They propose a Tree-oriented MapReduce framework for long-context reasoning . it leverages the hierarchical structure of long documents by constructing a DocTree .
Outcome: Experimental results show that ToM outperforms existing divide-and-conquer frameworks and RAGs . the proposed framework improves logical coherence and long-context reasoning on 70B+ LLMs compared to existing approaches .
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons (2026.acl-long)

Copied to clipboard

Challenge: Autoregressive (AR) models rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregression models.
Approach: They propose a framework that efficiently adapts autoregressive (AR) models to the diffusion paradigm.
Outcome: The proposed framework reduces training costs by orders of magnitude while maintaining state-of-the-art performance.
Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Long-video understanding is bottlenecked by the high cost of processing massive visual tokens.
Approach: They propose a decoupled framework for query-guided visual token pruning . their method reduces visual tokens by 90% and accelerates inference by 98% .
Outcome: The proposed framework reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average.
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model (2025.findings-naacl)

Copied to clipboard

Challenge: Existing general-domain visual language models lack ability of music notation understanding . Symbolic music is represented in two distinct forms: auditory music and symbolic music .
Approach: They propose to train a multimodal music notation model using a large-scale dataset . they use cross-modal alignment to train the model for music notations analysis .
Outcome: The proposed model improves on music understanding by training with a multimodal music notation model.
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities across various fields, but their widespread use is facing a severe and realistic challenge, which is their high demand for GPU memory.
Approach: They propose a KV cache reduction method which balances both shallow and deep layers by using an attention weight based eviction method and a codebook based replacement approach.
Outcome: The proposed method reduces the KV cache for shallower layers while preserving similar or even better model performance.
FSUIE: A Novel Fuzzy Span Mechanism for Universal Information Extraction (2023.acl-long)

Copied to clipboard

Challenge: Existing Universal Information Extraction models rely heavily on span boundaries in data during training, which does not reflect the reality of span annotation challenges.
Approach: They propose a framework that uses fuzzy spans to model various IE tasks . they propose generative Universal Information Extraction (UIE) to unify various ie tasks based on fuzzy span boundaries .
Outcome: The proposed framework improves on a series of main IE tasks with small amounts of data and training epochs.
VHASR: A Multimodal Speech Recognition System With Vision Hotwords (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that incorporate audio-related image information do not improve speech recognition performance.
Approach: They propose a novel approach utilizing audio-related image information and set up a multimodal speech recognition system that uses vision as hotwords to enhance the model’s speech recognition capability.
Outcome: The proposed model outperforms unimodal ASR model and achieves SOTA among existing image-based multimodal ASL models.
GoT-R1: Internalizing Graph-of-Thought via Structural Reinforcement for High-Density Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Chain-of-Thought reasoning suffers from an inherent mechanism flaw: linearity induces overthinking . emergence of Large Language Models (LLMs) has fundamentally redefined artificial intelligence .
Approach: They propose a framework that replaces verbose linear trajectories with high-density reasoning graphs.
Outcome: The proposed framework outperforms state-of-the-art models with reduced token overhead.
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to measure word segmentation only assess the language model's understanding of the overall meaning of sentences, lacking an evaluation of the language models' understanding capabilities at a fine-grained level.
Approach: They propose a framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) they employ current mainstream LLMs to perform word segmentations across multiple languages .
Outcome: The proposed method improves on existing methods and combines the advanced pattern recognition capabilities of Aho-Corasick automata with the deep insights of well-pretrained LLMs.
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction (2025.naacl-long)

Copied to clipboard

Challenge: Extractive UIEs can solve model explosion problems using a relatively small model . single-target instruction UIE enables the extraction of only one type of relation at a time .
Approach: They propose a model that assigns different relations to different levels for understanding and decision-making.
Outcome: Experiments show that LDNet outperforms state-of-the-art systems on 9 tasks, 33 datasets . LDnet outperformed state- of-the art systems on single-modal and multi-modal tasks .
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to prune redundant vision tokens struggle in shallow layers due to the lack of contextual information.
Approach: They propose a layer-wise contextualized visual token pruning method that uses a plug-and-play Pruning Module to prune redundant vision tokens.
Outcome: The proposed method outperforms training-free pruning methods under equal token budgets and surpasses training based methods with comparable supervision.
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.
Approach: They propose a paradigm called KV-Latent to reduce the KV cache footprint and improve inference speed by down-sampling the Key-Value vector dimensions into a latent space.
Outcome: The proposed paradigm reduces the KV Cache footprint and improves inference speed with a small amount of extra training, less than 1% of pre-training takes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations