Papers by Lidan Wang
SkipBERT: Efficient Inference with Shallow Layer Skipping (2022.acl-long)
Copied to clipboard
| Challenge: | Pre-trained language models have significant demands in computation and inference time, limiting their use in resource-constrained or latencysensitive applications. |
| Approach: | They propose to encode text chunks into independent representations and skip computation of shallow layers to accelerate inference. |
| Outcome: | The proposed approach can reduce latency by 65% without sacrificing performance. |
Draft
& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for accelerating Large Language Models have been criticized for their inference costs and inefficient decoding. |
| Approach: | They propose a self-speculative decoding approach for accelerating Large Language Models without an auxiliary model. |
| Outcome: | The proposed method achieves a speedup of up to 1.99 with no additional neural network training and no extra memory footprint. |
Open-Domain Question Answering with Pre-Constructed Question Spaces (2021.naacl-srw)
Copied to clipboard
| Challenge: | Open-domain question answering aims at locating answers to user-generated questions in massive collections of documents. |
| Approach: | They propose an algorithm with a novel reader-retriever design that differs from both families of algorithms. |
| Outcome: | The proposed algorithm outperforms retrieval-based methods with two large-scale datasets and is state-of-the-art. |
MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent work suggests a prefill-stage KV cache selection method to estimate KV importance from prefilling statistics. |
| Approach: | They propose a training-free, decode-aware and strictly prefill-only KV selection method that retains key-value caching for decoding . |
| Outcome: | The proposed method outperforms existing methods under tight cache budgets on multimodal benchmarks. |
Pyramid: A Layered Model for Nested Named Entity Recognition (2020.acl-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a fundamental NLP task. |
| Approach: | They propose a pyramid-like layered model for Nested Named Entity Recognition . token or text region embeddings are recursively inputted into L flat NER layers . |
| Outcome: | The proposed model achieves state-of-the-art F1 scores in nested NER on ACE-2004, ACE 2005, GENIA, and NNE. |
Understanding Points of Correspondence between Sentences for Abstractive Summarization (2020.acl-srw)
Copied to clipboard
| Challenge: | Using points of correspondence, fusion systems are difficult for abstractive summarizers because of their complexity. |
| Approach: | They propose to model points of correspondence between disparate sentences by combining documents, source and fusion sentences, and human annotations of points of correspondance between sentences. |
| Outcome: | The proposed model bridges the gap between coreference resolution and summarization by using human annotations of points of correspondence between sentences. |
Learning to Fuse Sentences with Transformers for Summarization (2020.emnlp-main)
Copied to clipboard
| Challenge: | Abstractive summarization systems that fuse sentences are not rewarded for correctly fusing sentences. |
| Approach: | They propose to leverage the knowledge of points of correspondence between sentences to enhance their ability to fuse sentences. |
| Outcome: | The proposed algorithms improve the ability of the proposed summarization systems to fuse sentences and show that they can fuse sentences in a way that retains the original meaning. |
MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety alignment methods, such as RLHF, fall into a Safety-Utility Trade-off, resulting in severe over-rejection of benign household instructions. |
| Approach: | They propose a meta-cognitive Critical Agent that evaluates peer debates using a structured argumentation framework derived from the Toulmin Model. |
| Outcome: | The proposed architecture outperforms existing systems in the SafeAware-VH benchmark. |
See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for video understanding suffer from autoregressive generation of tokens. |
| Approach: | They propose a training-free loosely SD framework for Video-LLMs that uses visual-relevant tokens to accurately pinpoint the latter. |
| Outcome: | The proposed framework boosts the accepted length and speedup ratio by 136% and 35% compared to SOTA training-free SD methods for Video-LLMs. |