Formal Verification of Mechanistic Interpretability Interventions (In Progress)
Developing methods to formally verify safety properties of activation-editing techniques in transformer models, tested on toy models, aimed at a top-tier conference submission. This is active, unpublished research — progress notes will be posted on the blog as the work continues.
Representative image to be added.
