Papers by Ethan Huang
Red Teaming Language Models with Language Models (2022.emnlp-main)
Copied to clipboard
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, Geoffrey Irving
| Challenge: | Prior work has found that language models (LMs) can harm users in hard-to-predict ways, and human annotation is expensive, limiting the number and diversity of test cases. |
| Approach: | They propose to generate test inputs using an LM itself, and use a classifier to detect harmful behavior on test input. |
| Outcome: | The proposed approach detects tens of thousands of offensive responses in a 280B parameter LM chatbot. |
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark (2025.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has not explored the mechanisms underlying this sensitivity. |
| Approach: | They propose a synthetic benchmark to evaluate Large Language Models’ reasoning robustness against systematically controlled irrelevant context (IC). |
| Outcome: | The proposed model improves in-distribution and out-of-disttribution scenarios while training with strong distractors. |