Papers by Zirui Li

22 papers
ClusterAttn: KV Cache Compression under Intrinsic Attention Clustering (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for sparse attention apply the same pattern across different attention heads and inputs, but fail to capture the intrinsic attention clustering in large language models.
Approach: They propose a training-free sparse attention method that provides an efficient prompt cache compression scheme under intrinsic attention clustering for efficient LLM inference.
Outcome: The proposed method reduces memory usage by 10%–65% and increases throughput by 2.6–4.8 times with no accuracy loss.
DocMMIR: A Framework for Document Multi-modal Information Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multi-modal information retrieval models lack a comprehensive exploration of document-level retrieval . existing models suffer from the absence of cross-domain datasets at this granularity.
Approach: They propose a multi-modal document retrieval framework to unify diverse document formats and domains with a comprehensive retrieval scenario.
Outcome: The proposed framework improves document retrieval performance on a large multimodal dataset.
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: specialized LLMs are often limited in domain-specific applications that require specialized knowledge.
Approach: They provide a comprehensive overview of four key methods to enhance large language models by integrating domain-specific knowledge.
Outcome: The proposed methods are categorized into four key approaches: dynamic knowledge injection, static knowledge embedding, modular adapters, and prompt optimization.
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models (2026.acl-long)

Copied to clipboard

Challenge: a recent study evaluated large audio-language models against jailbreak attacks . a new benchmark is being developed to evaluate LAM safety against jailbreaking attacks based on temporal and semantic nature of speech .
Approach: They propose a benchmark to evaluate LAM jailbreak vulnerabilities in adversarial audio prompts . they use a dataset of 1,495 adversarials to evaluate their performance .
Outcome: The proposed benchmark evaluates state-of-the-art LAMs against jailbreak attacks . it demonstrates that even small, semantically preserved perturbations can reduce safety .
Augmenting Operations Research with Auto-Formulation of Optimization Models From Problem Descriptions (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing systems for operations research use NLP to suggest formulations of optimization problems.
Approach: They propose an augmented intelligence system that can be used to simplify and enhance the modeling experience for operations research.
Outcome: The proposed system validates and edits the proposed formulations with a dataset of linear programming problems drawn from various application domains.
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches (2024.findings-emnlp)

Copied to clipboard

Challenge: Long context capability is a crucial competency for large language models as it mitigates the human struggle to digest long-form texts.
Approach: They propose to evaluate 10+ state-of-the-art approaches for long context-capable LLMs.
Outcome: The proposed methods are compared against 10+ state-of-the-art approaches across seven categories of long context tasks.
Recipe2Plan: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks focus on single task performance, ignoring multitask planning and execution efficiency.
Approach: They propose a benchmark framework based on real-world cooking scenarios . recipe2plan challenges agents to optimize cooking time through parallel task execution .
Outcome: The proposed benchmarks highlight the need for improved temporal awareness and global multitasking capabilities in large language models.
Controllable Contamination Detection for Reliable LLM Evaluation with Statistical Guarantees (2026.acl-long)

Copied to clipboard

Challenge: Existing training data detectors fail to detect clean samples from contaminated test sets . existing methods fail to identify clean samples due to black-box nature of LLMs .
Approach: They propose a framework that detects and filters contaminated evaluation data . they propose 'failure detection' to reduce the proportion of contaminated samples mistakenly retained .
Outcome: The proposed framework reduces false discovery rate (FDR) under valid FDR control while maintaining evaluation consistency.
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large audio-language models (LALMs) can exhibit a temporal smoothing bias . unified decoders can produce less specific audio-grounded outputs .
Approach: They propose a temporally blurred slow-path view that is re-encoded by a token-level logit update.
Outcome: Experiments on MMAU and AIR-Bench show consistent improvements on strong unified LALMs.
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision (2026.findings-acl)

Copied to clipboard

Challenge: Egocentric AI agents rely on pointing to resolve referential ambiguities in natural language commands.
Approach: They propose a question-answering benchmark to evaluate and enhance pointing reasoning in egocentric views.
Outcome: The proposed benchmark evaluates and enhances pointing reasoning in egocentric views.
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models have led to a growing interest in tool assisted LLMs . toolSandbox includes stateful tool execution, implicit state dependencies between tools .
Approach: a new tool-based evaluation tool is released to help LLMs evaluate their tool-use capabilities. a tool-driven evaluation tool includes stateful tool execution, implicit state dependencies between tools and a built-in user simulator.
Outcome: the toolSandbox evaluation benchmark shows that open source and proprietary models have a performance gap . the benchmarks show that even the most capable LLMs are challenged by state dependent tasks .
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting (2026.findings-acl)

Copied to clipboard

Challenge: FineState-Bench evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state.
Approach: They propose a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state.
Outcome: The proposed benchmark evaluates whether an agent can ground an instruction to the intended UI control and reach the exact target state.
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem (2025.findings-emnlp)

Copied to clipboard

Challenge: distributing LLMs without a proven track record like ‘meta-llama‘ or ‘qwen‘ rarely gains community traction.
Approach: They propose a simple, efficient, yet specific recipe for a backdoor LoRA to be injected into task-enhancing LoRAs and examine the mechanisms of such infections.
Outcome: The proposed model allows attackers to scale the distribution of compromised LoRAs with minimal effort by leveraging the rich pool of shared LoRA assets.
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection (2026.acl-long)

Copied to clipboard

Challenge: Personalized MGT detection remains largely underexplored due to personalization challenges . large language models (LLMs) can imitate personal writing styles, but they can generate fake news and misinformation.
Approach: They propose a benchmark to evaluate detector robustness under personalization . they attribute this limitation to a feature-inversion trap that flips the effect in personalized contexts .
Outcome: The proposed framework predicts detector robustness under personalization with an 85% correlation to actual results.
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for MU are limited by a lack of image diversity and coarse-grained unlearning targets.
Approach: They propose a benchmark to evaluate misinformation unlearning in MLLMs . OFFSIDE supports advanced unlearning targets such as fine-grained unlearning and visual rumor removal.
Outcome: OFFSIDE supports advanced unlearning targets, such as fine-grained unlearning and visual rumor removal.
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images (2026.acl-long)

Copied to clipboard

Challenge: AEGIS examines whether current models can effectively audit AI-generated images in academic papers.
Approach: They propose a holistic benchmark for forensic analysis of AI-Generated academic ImageS that reveals limitations in academic image forensics.
Outcome: AEGIS compared with existing benchmarks on seven academic categories and features key advances in forensic analysis.
Choosing Transfer Languages for Cross-Lingual Learning (P19-1)

Copied to clipboard

Challenge: Cross-lingual transfer is a useful tool for improving performance of natural language processing (NLP) on low-resource languages.
Approach: They propose to use cross-lingual transfer to improve accuracy of low-resource languages . they build models that consider features to perform prediction on such languages based on ranking problem .
Outcome: The proposed model predicts good transfer languages much better than baselines considering single features in isolation.
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on evaluating MLLMs’ pre-existing knowledge or perceptual understanding, often neglecting the critical capability of reasoning.
Approach: They propose a benchmark designed for visual clue-driven reasoning in daily scenarios that combines rigorous grounding in authentic daily activities and challenging query design that necessitates more than surface-level perception.
Outcome: The proposed benchmark identifies visual clues and their ability to provide robust reasoning in daily scenarios.
Re3: Relevance & Recency Retrieval for Mitigating Temporal Hallucination (2026.acl-long)

Copied to clipboard

Challenge: Existing retrievers suffer from temporal-semantic misalignment and outdated-document interference . Existing frameworks suffer from both temporal validity and outdated factual versions .
Approach: They propose a framework that mitigates temporal hallucinations by embedding heterogeneous temporal signals into the semantic space to ensure retrieval fidelity.
Outcome: Experiments show that Re3 outperforms baselines by 9.7% in generation accuracy . the framework outperformed strongest baselines on challenging dynamic tasks .
Example Quality Matters: Multi-Aspects Example Augmentation for Private Library Programming (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to code generation fail to consider the quality of retrieved examples.
Approach: They propose a retrieval-augmented generation method that combines existing API examples to improve complexity and readability.
Outcome: The proposed method achieves up to 22% accuracy improvement over baseline methods.
ModularMoE: Fast LLM Customization with Parameter-Sharing Mixture-of-Experts for Low-Resource Settings (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models impose significant computational and storage burdens on personal devices . existing customization approaches incur excessive computational costs or lead to suboptimal performance .
Approach: They propose a training framework that converts pre-trained LLMs into parameter-sharing MoE models for lightweight deployment.
Outcome: The proposed training framework outperforms state-of-the-art training frameworks at the same sparsity level while delivering up to 2.71 inference speedup.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations