Papers by Mahmoud Khalil
The Unintended Trade-off of AI Alignment: Balancing Hallucination Mitigation and Safety in LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Hallucination in large language models has been studied, but a side effect remains unrecognized . a new study examines the trade-off between truthfulness and safety alignment . |
| Approach: | They propose a method that disentangles hallucination from hallucinian features using sparse autoencoders. |
| Outcome: | The proposed method preserves refusal behavior and task utility while maintaining safety alignment. |
An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels (2022.acl-long)
Copied to clipboard
Taylor Sorensen, Joshua Robinson, Christopher Rytting, Alexander Shaw, Kyle Rogers, Alexia Delorey, Mahmoud Khalil, Nancy Fulda, David Wingate
| Challenge: | Existing prompt engineering methods require labeled data and access to model parameters . a new method for selecting prompt templates without labeles and without direct access to the model is needed. |
| Approach: | They propose a method for selecting prompt templates without labeled examples and without direct access to the model. |
| Outcome: | The proposed method performs at almost oracle levels, without labels, on 7 datasets representing 7 different NLP tasks. |