Papers by Hanxu Hu
CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models (2024.findings-naacl)
Copied to clipboard
Wenhong Zhu, Hongkun Hao, Zhiwei He, Yun-Ze Song, Jiao Yueyang, Yumeng Zhang, Hanxu Hu, Yiran Wei, Rui Wang, Hongyuan Lu
| Challenge: | Existing methods to evaluate large language models are prone to data contamination. |
| Approach: | They propose a method which parses contaminated data and back-translates it into a candidate set. |
| Outcome: | The proposed method reduces data contamination and evaluates the LLMs more cleanly. |
Improving User Controlled Table-To-Text Generation Robustness (2023.findings-eacl)
Copied to clipboard
| Challenge: | In experiments, models perform well on test sets coming from the same distribution as the train data but their performance drops when evaluated on realistic noisy user inputs. |
| Approach: | They propose a user controlled table-to-text generation task where users explore the content in a table by selecting cells and reading a natural language description thereof. |
| Outcome: | The proposed model gains 4.85 BLEU points on user noisy test cases and 1.4 on clean test cases. |
Fine-Tuning Large Language Models with Sequential Instructions (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing instruction-tuned models struggle to adhere to a query with multiple intentions, which impairs their performance when the completion of several tasks is demanded by a single command. |
| Approach: | They develop an automatic process that turns existing data into diverse and complex task chains and a new benchmark to evaluate a model’s ability to follow all the instructions in a sequence. |
| Outcome: | The proposed model can follow instructions better and deliver higher results in coding, maths, and open-ended generation. |
LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a problem of data contamination is now almost inevitable during the development of large language models, with the training data often integrating evaluation benchmarks even unintentionally. |
| Approach: | They propose a framework to restore model performance prior to data contamination on potentially leaked datasets by using contamination detection and disruption operation. |
| Outcome: | The proposed framework restores model performance prior to contamination on potentially leaked datasets. |
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multilingual benchmarks focus primarily on language understanding tasks. |
| Approach: | They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages. |
| Outcome: | Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve. |
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Document-level machine translations have paved the way for truly simple document-level translation, but challenges such as omission errors remain. |
| Approach: | They propose a method for document-level machine translation that leverages previous contexts in a multi-turn conversational manner by decomposing documents into segments and iteratively translating them while maintaining previous turns. |
| Outcome: | The proposed method outperforms translations of entire documents in a single turn and translations independently according to multiple automatic metrics in representative LLMs. |