Papers by Chenxi Gu
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models can generate highquality, human-like content, but they also pose risks such as infringement of proprietary interests, misuse of outputs, and spread harmful misinformation. |
| Approach: | They propose a method that partitions the vocabulary into two logit-balanced subsets and lifts the lower bound of watermark strength for each token prediction. |
| Outcome: | The proposed method lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability. |
Word Form Matters: LLMs’ Semantic Reconstruction under Typoglycemia (2025.findings-acl)
Copied to clipboard
| Challenge: | Typoglycemia is a phenomenon where people can read words even when the middle letters of the words are scrambled. |
| Approach: | They propose a reliable metric to quantify the degree of semantic reconstruction and validate its effectiveness. |
| Outcome: | The proposed metric quantifies the degree of semantic reconstruction and validates its effectiveness. |
Leveraging Similar Users for Personalized Language Modeling with Limited Data (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work suggests that personalized models are more accurate for individual users than one-size-fits-all solutions. |
| Approach: | They propose a model trained on users that are similar to a new user to find similarity between new and existing users. |
| Outcome: | The proposed model can predict what a user will write when they join a platform and not enough text is available. |
Watermarking PLMs on Classification Tasks by Combining Contrastive Learning with Weight Perturbation (2023.findings-emnlp)
Copied to clipboard
Chenxi Gu, Xiaoqing Zheng, Jianhan Xu, Muling Wu, Cenyuan Zhang, Chengsong Huang, Hua Cai, Xuanjing Huang
| Challenge: | Large pre-trained language models (PLMs) are highly valuable intellectual property due to their expensive training costs. |
| Approach: | They propose to embed backdoors that can be triggered by specific inputs into models by model watermarking. |
| Outcome: | The proposed method can be used to protect the intellectual property of large pre-trained language models without knowledge about downstream tasks. |
Unified Hallucination Detection for Multimodal Large Language Models (2024.acl-long)
Copied to clipboard
Xiang Chen, Chenxi Wang, Yida Xue, Ningyu Zhang, Xiaoyan Yang, Qiang Li, Yue Shen, Lei Liang, Jinjie Gu, Huajun Chen
| Challenge: | despite significant strides in multimodal tasks, MLLMs are plagued by the critical issue of hallucination. |
| Approach: | They propose a meta-evaluation benchmark to facilitate evaluation of advancements in hallucination detection methods. |
| Outcome: | The proposed framework validates hallucinations robustly and provides strategic insights . MHaluBench is a meta-evaluation benchmark designed to facilitate evaluation . |