Papers with DMoERM
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling (2024.findings-acl)
Copied to clipboard
| Challenge: | Using a reward model (RM) to improve the effectiveness of large language models, there are two challenges in training. |
| Approach: | They propose a reward model (RM) that is a proxy of human preferences and assigns scores to the outputs of the large language model (LLM) a human annotation consistency rate of 60% to 75% is causing training data to contain a lot of noise. |
| Outcome: | The proposed model outperforms state-of-the-art ensemble methods and mitigates the overoptimization problem. |