Papers by Martin Mueller
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)
Copied to clipboard
| Challenge: | a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE. |
| Approach: | They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation. |
| Outcome: | The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI. |
CRISP: Persistent Concept Unlearning via Sparse Autoencoders (2026.acl-long)
Copied to clipboard
| Challenge: | Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features, but most SAE-based methods operate at inference time, which does not create persistent changes in the model’s parameters. |
| Approach: | They propose a parameter-efficient method for persistent concept unlearning using SAEs that automatically identifies salient SAE features across multiple layers and suppresses their activations. |
| Outcome: | The proposed method outperforms previous methods on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. |
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have shown remarkable capabilities, particularly in English, but for less prevalent languages, performance can be significantly lower, making additional adaptation paramount. |
| Approach: | They propose a new adaptation method based on iteratively merging multiple models fine-tuned on a subset of available training data that reduces forgetting while maintaining learning on the target domain. |
| Outcome: | The proposed method outperforms LLAMA-3-8B-based models in German and German while maintaining learning on the target domain. |
Discovering Language Model Behaviors with Model-Written Evaluations (2023.findings-acl)
Copied to clipboard
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, Jared Kaplan
| Challenge: | Prior work creates evaluations with crowdwork or existing data sources, which are not always available. |
| Approach: | They generate evaluations automatically with language models (LMs) using crowdwork or existing data sources to find out how they behave . |
| Outcome: | The results show that large LMs repeat back a dialog user’s preferred answer and express greater desire to pursue concerning goals like resource acquisition and goal preservation. |