Papers by Shanshan Wu
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Detecting AI-generated poetry is difficult due to distinctive characteristics of modern Chinese poetry. |
| Approach: | They propose a benchmark for detecting AI-generated modern Chinese poetry . they use a high-quality dataset and systematic performance assessments . |
| Outcome: | The proposed benchmark is based on a high-quality dataset of 800 poems written by six professional poets and 41,600 poems generated by four mainstream LLMs. |
Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications (2025.acl-industry)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have achieved impressive performance on many language tasks. |
| Approach: | They synthesize a high-quality dataset of error correction pairs to evaluate and improve LLMs for mobile applications by reweighting the sample. |
| Outcome: | The proposed model improves on offline evaluation and live A/B testing, given the LLM performance on offline data and scores from a small privacy-preserving on-device language model. |
Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent research indicates that Large Reasoning Models suffer from a strategic bottleneck at reasoning path planning. |
| Approach: | They propose a framework that reformulates reasoning as a dynamic search for the optimal thinking strategy. |
| Outcome: | The proposed framework improves accuracy and computational cost while reducing generation length by over 22%. |
Training-free LLM Merging for Multi-task Learning (2025.acl-long)
Copied to clipboard
Zichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang, Shanshan Ye, Hongzhi Yin, Yi Chang, Yefeng Zheng, Xiangyu Zhao
| Challenge: | Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing tasks. |
| Approach: | They propose a training-free method for unifying different specialized LLMs into a single model using model-wise and layer-wise pruning and scaling. |
| Outcome: | The proposed method outperforms existing merging techniques and surpasses models fine-tuned on combined datasets in most scenarios. |
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a lack of comprehensive benchmarks for Routing large language models has hindered the development of routers. |
| Approach: | They propose a router-based benchmark to evaluate Routing large language models . the benchmark includes performance records for 12 popular LLM evaluations . |
| Outcome: | The proposed model-level scaling up phenomenon can surpass the best single model in the pool and many existing strong LLMs. |
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing LCU benchmarks for large language models often result in prohibitively high evaluation costs . existing benchmarks exhibit significant redundancy, which means inefficiency in evaluation . |
| Approach: | They propose a data compression method tailored for long-text data with sparse information characteristics. |
| Outcome: | The proposed method reduces evaluation costs to 4.5% of the long-text benchmark LongBench . the proposed method is based on a long-term LCU benchmark with sparse information characteristics . |
Weight-Inherited Distillation for Task-Agnostic BERT Compression (2024.findings-naacl)
Copied to clipboard
| Challenge: | Knowledge Distillation (KD) is a predominant approach for BERT compression. |
| Approach: | They propose a weight-inherited distillation method which directly transfers knowledge from the teacher to a compact student model by inheriting the weights. |
| Outcome: | The proposed method outperforms state-of-the-art KD-based methods on GLUE and SQUAD benchmarks. |