Papers by Ali Khalil

2 papers
The Unintended Trade-off of AI Alignment: Balancing Hallucination Mitigation and Safety in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Hallucination in large language models has been studied, but a side effect remains unrecognized . a new study examines the trade-off between truthfulness and safety alignment .
Approach: They propose a method that disentangles hallucination from hallucinian features using sparse autoencoders.
Outcome: The proposed method preserves refusal behavior and task utility while maintaining safety alignment.
Efficient Citer: Tuning Large Language Models for Enhanced Answer Quality and Verification (2024.findings-naacl)

Copied to clipboard

Challenge: Existing models with explicit citations lack the ability to verify information generated by these models.
Approach: They construct a citation training dataset and fine-tune two models to address the challenge of explicit citations efficiently.
Outcome: The proposed models surpass ChatGPT and exhibit exceptional out-of-domain generalization in both human and automatic evaluation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations