Papers by Koh Hng
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models (2024.emnlp-demo)
Copied to clipboard
Prannaya Gupta, Le Yau, Hao Low, I-Shiang Lee, Hugo Lim, Yu Teoh, Koh Hng, Dar Liew, Rishabh Bhardwaj, Rajat Bhardwaj, Soujanya Poria
| Challenge: | Potential harms include training data leakage, biases in responses and decision-making, and unauthorized use for purposes such as terrorism and the generation of sexually explicit content. |
| Approach: | WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models. |
| Outcome: | The framework supports both LLM and judge benchmarking and incorporates custom mutators to test safety against various text-style mutations such as future tense and paraphrasing. |