Papers by Yongchan Kwon
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior studies assess instruction adherence in the model’s main responses, but it is also critical for large reasoning models to follow user instructions throughout their reasoning process. |
| Approach: | They propose a systematic benchmark for assessing reasoning instruction following to assess the model's adherence to instructions. |
| Outcome: | The proposed benchmark reduces the risk of undesirable shortcuts, hallucinations, or reward hacking within reasoning traces. |
Understanding Impact of Human Feedback via Influence Functions (2025.acl-long)
Copied to clipboard
| Challenge: | In reinforcement learning from human feedback, human feedback can be noisy, inconsistent or biased . this variability can lead to misaligned reward signals, potentially causing unintended side effects . |
| Approach: | They propose an approximation method that measures the impact of human feedback on the performance of reward models. |
| Outcome: | The proposed method detects common labeler biases in human feedback datasets and guides labelers in refining their strategies to better align with expert feedback. |