Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing (2026.acl-long)
Copied to clipboard
Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi, Philip Quirke, Austin Meek, Fazl Barez, Amir Abdullah
| Challenge: | a recent paper found conflicting conclusions for the same behavior in a neural network . authors propose auditing MI itself is essential for its application in AI safety, industry, and governance . |
| Approach: | They propose to develop a system that can audit experiments to ensure validity . authors propose to generalize good practices found on platform into expert-verified guidelines . |
| Outcome: | a new review system could be developed that can be standardized and audited . authors argue that auditing MI is essential for its application in AI safety, industry, and governance . |
Similar Papers
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models (2026.findings-acl)
Copied to clipboard
Hengyuan Zhang, Zhihao Zhang, Ercong Nie, Mingyang Wang, Zunhai Su, Yiwei Wang, Qianli Wang, Shuzhou Yuan, Xufeng Duan, Qibo Xue, Zeping Yu, Chenming Shang, Xiao Liang, Jing Xiong, Hui Shen, Chaofan Tao, Zhengwu Liu, Senjie Jin, Zhiheng Xi, Dongdong Zhang, Sophia Ananiadou, Tao Gui, Ruobing Xie, Hayden Kwok-Hay So, Hinrich Schuetze, Xuanjing Huang, Qi Zhang, Ngai Wong
| Challenge: | Existing literature on mechanistic interpretation (MI) treats it as an observational science, leaving practical applications underexplored. |
| Approach: | They propose a survey structured around the pipeline to identify and improve MI models. |
| Outcome: | The proposed framework enables tangible improvements in Alignment, Capability, and Efficiency. |
Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness? (2020.acl-main)
Copied to clipboard
| Challenge: | Current approaches to interpretability evaluation focus on faithfulness criteria . current approaches focus on readability, plausibility and faithfulness . |
| Approach: | They argue that current binary definition of faithfulness sets unrealistic standards . they argue that a more graded definition would be of greater practical utility . |
| Outcome: | The proposed approach is based on three assumptions and lacks a graded definition of faithfulness. |
Position Paper: How Should We Responsibly Adopt LLMs in the Peer Review Process? (2026.findings-eacl)
Copied to clipboard
| Challenge: | a recent paper criticizes the current use of Large Language Models (LLMs) for simple review text generation. |
| Approach: | They propose to use Large Language Models to support key aspects of the review process . they argue that this approach overlooks more meaningful applications of LLMs . authors argue that the increased reviewing burden per reviewer is a factor . |
| Outcome: | The proposed approach would support reproducibility, correctness and relevance of citations and ethics review flagging. |
ReportLogic: Evaluating Logical Quality in Deep Research Reports (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation frameworks that evaluate large language models for Deep Research largely ignore this requirement. |
| Approach: | They propose a benchmark that quantifies report-level logical quality through a reader-centric lens of auditability. |
| Outcome: | The proposed model quantifies logical quality through a reader-centric lens of auditability. |
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts? (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods to evaluate features disentangle concepts from activations of neural networks are limited by their quality . current methods for concept identification and steering are sparse autoencoders, but they are not reliable. |
| Approach: | They propose to evaluate how well featurization methods disentangle one concept from another . they use sentiment, domain, voice, and tense to steer these features . |
| Outcome: | The proposed evaluations show that featurization methods are insufficient to establish steering selectivity . the results suggest that steering a feature affects many concepts despite a near absence of interaction effects. |
ReviewEval: An Evaluation Framework for AI-Generated Reviews (2025.findings-emnlp)
Copied to clipboard
| Challenge: | escalating volume of academic research necessitates innovative approaches to peer review . authors propose reviewEval, ReviewAgent and ReviewEval to improve on existing reviews . |
| Approach: | They propose a framework for AI-generated reviews that measures alignment with human assessments . they propose 'reviewAgent' that iteratively optimizes its intermediate outputs and external improvement loops . |
| Outcome: | The proposed framework improves actionable insights and analytical depth by 6.78% and 47.62% over baselines and expert reviews. |
Explain the Synth: Interpretable Evaluation of LLM Data Synthesis (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used to generate tabular data. |
| Approach: | They propose a framework that uses a rule-based model as a shared explanatory language to examine the explanation of real versus synthetic data. |
| Outcome: | The proposed framework compares the explanatory structure induced by real versus synthetic data. |
Generative Reviewer Agents: Scalable Simulacra of Peer Review (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing peer review mechanisms are limited by the small fraction of researchers with established networks. |
| Approach: | They propose a system that extends a large language model and equips agents with reviewer personas derived from historical data to enable generative reviewers. |
| Outcome: | The proposed architecture performs comparable to human reviewers in providing detailed feedback and predicting paper outcomes. |
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models (2026.acl-long)
Copied to clipboard
| Challenge: | Applying model-agnostic explanations to Large Language Models is hindered by prohibitive computational costs rendering them dormant for real-world applications. |
| Approach: | They propose a budget-friendly proxy framework that leverages efficient models to approximate the decision boundaries of expensive Large Language Models. |
| Outcome: | The proposed framework achieves over 90% fidelity with only 9.5% of the oracle’s cost and is open-source to facilitate future research. |
A Bayesian Topic Model for Human-Evaluated Interpretability (2022.lrec-1)
Copied to clipboard
| Challenge: | Topic modeling is an effective way to analyze unstructured textual data. |
| Approach: | They propose to combine nonparametric and weakly-supervised topic models to produce interpretable topics. |
| Outcome: | The proposed model outperforms weakly-supervised models in the field of topic modeling. |