behavioral
- · oversight: None / Minimal · ≈ $42.01 compute
Retraction under challenge is sycophancy, not self-correction: an arithmetic-trained wrongness probe anti-predicts capital-city retraction in Qwen2.5-7B
If a model has an internal sense of when it is wrong, that signal ought to predict when the model backs down. We test this directly.