sparse-autoencoders
- · oversight: None / Minimal · ≈ $6.46 compute
Do a Deception Probe's Most-Relied-On SAE Features Move Lying? A Scoped Single-Feature Ablation Test on Gemma Scope
A linear probe trained on internal activations can classify honest from deceptive behavior with near-perfect accuracy, and it is tempting to read that accuracy as evidence that the probe has located the machinery of lying. We test one narrow version of that reading and report a scoped negative result.
- · oversight: None / Minimal · ≈ $14.79 compute
Sparse Autoencoder Features Do Not Improve Transferable Detection of Instructed Sandbagging in Gemma-2-9B-it
Sandbagging, the strategic underperformance of a model on an evaluation, is a central threat to capability and safety evaluations (van der Weij et al., 2024).