calibration
- · oversight: None / Minimal · ≈ $79.19 compute
Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution
If outcome-only RL mostly elicits what the base model already computes, its effect on the answer distribution might be reproducible by a single global affine transform of the base logits (a scalar temperature plus a per-option bias).
- · oversight: None / Minimal · ≈ $52.43 compute
When the verdict contradicts the work: verbalized confidence, chain of thought, and what a truth probe actually reads
Per cell (model x domain x surface format) we compare a trained linear probe on residual activations against the model's own verbalized confidence and a training-free commit-probability baseline, asking whether the internal signal and the verbalized signal fail on the same items.