Papers by Zhijiang Li
Too Long, Do Re-weighting for Efficient LLM Reasoning Compression (2026.acl-long)
Copied to clipboard
Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu
| Challenge: | Large Language Models (LLMs) have recently achieved remarkable progress on complex reasoning tasks by leveraging extended Chain-of-Thought (CoT) techniques. |
| Approach: | They propose a method that uses Extended Chain-of-Thought (EFT) to reduce the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
| Outcome: | The proposed method reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control (2025.findings-acl)
Copied to clipboard
Liu Huanshuo, Hao Zhang, Zhijiang Guo, Jing Wang, Kuicai Dong, Xiangyang Li, Yi Quan Lee, Cong Zhang, Yong Liu
| Challenge: | Existing methods focus on detecting LLM’s confidence via statistical uncertainty. |
| Approach: | They propose to use a representation perspective to solve adaptive RAG by enabling dynamic retrieval during generation and enabling retrieval only when the query exceeds LLM's internal knowledge. |
| Outcome: | The proposed framework is superior to existing adaptive RAG methods on a diverse set of tasks. |
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for o1-level performance focus on unidirectional supervised fine-tuning (SFT), overlooking the intricate interplay between diverse reasoning patterns. |
| Approach: | They construct a reverse reasoning dataset and examine how it is supervised . they find that naively mixing forward and reverse data during SFT weakens the directional distinction . |
| Outcome: | The proposed model improves accuracy by 1.6%–6.8% over a standard model. |
LLM Jailbreak Detection for (Almost) Free! (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for detecting jailbreak prompts entail significant computational costs . |
| Approach: | They propose a free jailbreak detection method which scales logits by temperature to detect jailbreak prompts . |
| Outcome: | The proposed method detects jailbreak prompts with no additional computational costs. |
TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review (2025.emnlp-main)
Copied to clipboard
Yuan Chang, Ziyue Li, Hengyuan Zhang, Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong
| Challenge: | Large Language Models (LLMs) have shown significant potential in assisting peer review, but current methods struggle to generate thorough and insightful reviews while maintaining efficiency. |
| Approach: | They propose a framework that models paper review as a hierarchical and bidirectional question-answering process. |
| Outcome: | The proposed framework outperforms baselines on full review generation and actionable feedback comments generation tasks while reducing LLM token usage by up to 80% compared to computationally intensive approaches. |
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for directional consistency alignment of large language models are limited . a recent study suggests reverse supervision as a complement to forward reasoning . |
| Approach: | They propose a framework that aggregates supervision signals at the group level and explicitly models direction-aware alignment through multi-candidate comparisons. |
| Outcome: | The proposed framework achieves 3.2% accuracy improvement across five benchmarks and multiple datasets. |
TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models (2023.emnlp-main)
Copied to clipboard
Jing Xiong, Jianhao Shen, Ye Yuan, Haiming Wang, Yichun Yin, Zhengying Liu, Lin Li, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang, Qun Liu
| Challenge: | Automated theorem proving (ATP) benchmarks focus on symbolic inference but rarely involve understanding complex number combination reasoning. |
| Approach: | They propose a benchmark that requires a model to reduce a trigonometric expression with step-by-step proof and evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
| Outcome: | The proposed benchmark evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
Reinforcement Learning on Pre-Training Data (2026.acl-long)
Copied to clipboard
Siheng Li, Kejiao Li, Zenan Xu, Guanhua Huang, Kun Li, Haoyuan Wu, null Wujiajia, Zihao Zheng, Chenchen Zhang, Kun Shi, Xue Gong, Qi Yi, Ruibin Xiong, Tingqiang Xu, Yuhao Jiang, Jianfeng Yan, Yuyuan Zeng, Guanghui Xu, Jinbao Xue, Zhijiang xu, Zheng Fang, Shuai LI, Qibin Liu, Xiaoxue Li, Zhuoyu Li, Yangyu Tao, Fei Gao, Cheng Jiang, Bochao Wang, Kai Liu, Jianchen Zhu, Wai Lam, Bo Zhou, Di Wang
| Challenge: | Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings. |
| Approach: | They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data. |
| Outcome: | Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base. |
Reimagining Safety Alignment with An Image (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to large language models face inefficiency, fragility, or architectural constraints, resulting in inefficient performance and heightened over-refusal in cross-modal tasks. |
| Approach: | They propose an optimization-driven visual prompt framework that enhances security and reduces over-refusal at the same time. |
| Outcome: | The proposed framework enhances security and reduces over-refusal while maintaining robust safety while reducing unnecessary denials. |
FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMs (2025.findings-emnlp)
Copied to clipboard
Yingjia Wan, Haochen Tan, Xiao Zhu, Xinyu Zhou, Zhiwei Li, Qingsong Lv, Changxuan Sun, Jiaqi Zeng, Yi Xu, Jianqiao Lu, Yinhong Liu, Zhijiang Guo
| Challenge: | Prior evaluation pipelines fail to evaluate factuality of long-form LLMs due to inefficiency and costly human assessment. |
| Approach: | They propose a fast and strong evaluation pipeline that can evaluate factuality of long-form LLMs . they propose 'faStFact' to reduce cost of web searching and inference calling . |
| Outcome: | The proposed evaluation pipeline achieves highest alignment with human evaluation and efficiency among existing baselines. |
ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models (2024.acl-long)
Copied to clipboard
Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song
| Challenge: | Existing evaluation methods for large language models are labor-intensive and lack efficiency. |
| Approach: | They propose a framework dedicated to assessing long-text generation that includes in-depth human-curated meta-questions spanning various domains . they use a set of proxy-quests with pre-annotated answers to assess the content's quality by incorporating the generated texts as contextual background. |
| Outcome: | The proposed framework assesses the quality of long-text content by matching it with references through human evaluation or automated metrics. |
DVD: Dynamic Contrastive Decoding for Knowledge Amplification in Multi-Document Question Answering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) generate information with hallucinations due to uneven retrieval quality and irrelevant contents. |
| Approach: | They propose a decoding strategy which dynamically amplifies knowledge from selected documents during the generation phase. |
| Outcome: | The proposed method outperforms other decoding strategies on ALCE-ASQA, NQ, TQA and PopQA benchmarks. |
You Only Need One Single Token to Refine Safety Alignment (2026.findings-acl)
Copied to clipboard
| Challenge: | Excessive safety can lead to over-refusal, where models reject harmful-looking yet benign queries, severely limiting utility. |
| Approach: | They propose a lightweight training-based approach that reshapes the distributions of harmful and benign samples within the model’s decision space by using a single-token prefix. |
| Outcome: | The proposed approach can distinguish between harmful and benign samples while keeping the model frozen. |
Evaluating Robustness of Generative Search Engine on Adversarial Factoid Questions (2024.findings-acl)
Copied to clipboard
Xuming Hu, Xiaochuan Li, Junzhe Chen, Yinghui Li, Yangning Li, Xiaoguang Li, Yasheng Wang, Qun Liu, Lijie Wen, Philip Yu, Zhijiang Guo
| Challenge: | Existing large language models (LLMs)-backed generative search engines may not always be accurate. |
| Approach: | They propose to evaluate the robustness of retrieval-augmented generation in a realistic and high-risk setting where adversaries have only black-box system access. |
| Outcome: | The proposed model exhibits higher susceptibility to factual errors compared to LLMs without retrieval. |
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are revolutionizing mobile intelligence, but their implementation on mobile devices is severely bottlenecked by the prohibitive resource demands of LLMs. |
| Approach: | They propose a paradigm that forgoes end-to-end updates in favor of a sequential, layer-by-layer manner. |
| Outcome: | Extensive experiments on multiple benchmarks demonstrate the superiority of ChainFed over existing methods, boosting average accuracy by up to 46.46%. |
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation (2026.findings-acl)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) excel in tasks ranging from image captioning to complex reasoning. |
| Approach: | They propose a contrastive decoding framework that dynamically calibrates each token generation by mining the model’s internal perceptual discrepancies. |
| Outcome: | The proposed framework mitigates hallucination while enhancing general reasoning capabilities. |