Papers by Chenxi Gu

5 papers
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking (2026.acl-long)

Copied to clipboard

Challenge: Large language models can generate highquality, human-like content, but they also pose risks such as infringement of proprietary interests, misuse of outputs, and spread harmful misinformation.
Approach: They propose a method that partitions the vocabulary into two logit-balanced subsets and lifts the lower bound of watermark strength for each token prediction.
Outcome: The proposed method lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability.
Word Form Matters: LLMs’ Semantic Reconstruction under Typoglycemia (2025.findings-acl)

Copied to clipboard

Challenge: Typoglycemia is a phenomenon where people can read words even when the middle letters of the words are scrambled.
Approach: They propose a reliable metric to quantify the degree of semantic reconstruction and validate its effectiveness.
Outcome: The proposed metric quantifies the degree of semantic reconstruction and validates its effectiveness.
Leveraging Similar Users for Personalized Language Modeling with Limited Data (2022.acl-long)

Copied to clipboard

Challenge: Recent work suggests that personalized models are more accurate for individual users than one-size-fits-all solutions.
Approach: They propose a model trained on users that are similar to a new user to find similarity between new and existing users.
Outcome: The proposed model can predict what a user will write when they join a platform and not enough text is available.
Watermarking PLMs on Classification Tasks by Combining Contrastive Learning with Weight Perturbation (2023.findings-emnlp)

Copied to clipboard

Challenge: Large pre-trained language models (PLMs) are highly valuable intellectual property due to their expensive training costs.
Approach: They propose to embed backdoors that can be triggered by specific inputs into models by model watermarking.
Outcome: The proposed method can be used to protect the intellectual property of large pre-trained language models without knowledge about downstream tasks.
Unified Hallucination Detection for Multimodal Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: despite significant strides in multimodal tasks, MLLMs are plagued by the critical issue of hallucination.
Approach: They propose a meta-evaluation benchmark to facilitate evaluation of advancements in hallucination detection methods.
Outcome: The proposed framework validates hallucinations robustly and provides strategic insights . MHaluBench is a meta-evaluation benchmark designed to facilitate evaluation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations