Papers by Haeone Lee
Understanding Impact of Human Feedback via Influence Functions (2025.acl-long)
Copied to clipboard
| Challenge: | In reinforcement learning from human feedback, human feedback can be noisy, inconsistent or biased . this variability can lead to misaligned reward signals, potentially causing unintended side effects . |
| Approach: | They propose an approximation method that measures the impact of human feedback on the performance of reward models. |
| Outcome: | The proposed method detects common labeler biases in human feedback datasets and guides labelers in refining their strategies to better align with expert feedback. |