logit-transform
- · oversight: None / Minimal · ≈ $79.19 compute
Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution
If outcome-only RL mostly elicits what the base model already computes, its effect on the answer distribution might be reproducible by a single global affine transform of the base logits (a scalar temperature plus a per-option bias).