Papers by Sumit Yadav
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering (2026.acl-long)
Copied to clipboard
| Challenge: | Current safety alignment methods fail to identify intended benign task before refusing to respond. |
| Approach: | They propose a method that uses inference-time trajectory-shifting to guide model behavior . they show that LLMs persist in refusing inputs containing harmful content . |
| Outcome: | The proposed approach reduces over-refusals with minimal impact on utility. |