Papers by Sadat Shahriar
MEAV: Model Editing with Alignment Vectors for inference time LLM alignment in single and multidomain preference spectrum (2026.findings-acl)
Copied to clipboard
Sadat Shahriar, Zheng Qi, Nikolaos Pappas, Srikanth Doss, Kishaloy Halder, Monica Sunkara, Manuel Mager, Yassine Benajiba
| Challenge: | Existing training-time alignment methods require full retraining when a change is needed. |
| Approach: | They propose an inference-time model-editing-based alignment method that learns encoded representations of preference dimensions and allows dynamic adjusting of the model behavior. |
| Outcome: | The proposed method can be used to align large language models to human preferences . it reduces the cost of inference by half compared to the prompt engineering approach . |