Papers by Chaojie Wang
Removing Prompt-template Bias in Reinforcement Learning from Human Feedback (2025.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) has shown promise for enhancing pre-trained large language models to generate responses that align with human preferences and societal values. |
| Approach: | They propose a method to estimate prompt-template bias term during reward modeling and use it to calibrate reward scores. |
| Outcome: | The proposed method can be flexibly combined with existing algorithms of removing length bias, leading to a further improvement in the aspect of enhancing the quality of generated responses. |
Friendly Topic Assistant for Transformer Based Abstractive Summarization (2020.emnlp-main)
Copied to clipboard
| Challenge: | Abstractive document summarization is a comprehensive task in natural language processing. |
| Approach: | They propose a topic assistant that rearranges and learns document semantics . they propose TA that is compatible with Transformer-based models and user-friendly . |
| Outcome: | The proposed model is compatible with Transformer-based models and user-friendly. |
EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering (2021.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that data diversity affects the performance of LMs if we train a single LM over the entire dataset. |
| Approach: | They propose an autoencoding topic model with a mixture prior to perform clustering for the data. |
| Outcome: | The proposed model can learn knowledge from different samples while extracting cluster-specific features. |