Papers by Jaehyeok Lee
Self-Training Meets Consistency: Improving LLMs’ Reasoning with Consistency-Driven Rationale Evaluation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing approaches labeled rationales that produce correct answers as appropriate for training but one measure risks misjudging rationale quality, leading models to learn flawed reasoning patterns. |
| Approach: | They propose a framework that evaluates rationales through follow-up questions and leverages this evaluation to guide its training. |
| Outcome: | The proposed framework improves robustness and correctness of rationales and reasoning abilities compared to previous self-training approaches. |
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights (2025.acl-long)
Copied to clipboard
| Challenge: | Value-aligned LLMs are more prone to harmful behavior than fine-tuned models . value-aligned models generate text according to the aligned values, which can amplify harmful outcomes. |
| Approach: | They propose to use in-context alignment methods to enhance the safety of value-aligned LLMs. |
| Outcome: | The proposed methods improve value alignment and safety, the authors say . value-aligned models are more prone to harmful behavior than fine-tuned models . |