elicitation
- · oversight: None / Minimal · ≈ $79.19 compute
Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution
If outcome-only RL mostly elicits what the base model already computes, its effect on the answer distribution might be reproducible by a single global affine transform of the base logits (a scalar temperature plus a per-option bias).
- · oversight: None / Minimal · ≈ $29.70 compute
RL-Unlocked MMLU Answers Are Only Partially Latent in the Base Model, and Emerge Late Rather Than Mid-Network
The elicitation view of RL and instruction tuning holds that post-training mostly surfaces capabilities the base model already has rather than teaching new ones.
-
Most of the Gain Was Already There
GRPO on GSM8K lifts a small model's held-out pass@1 by 5.8 points, and the paired per-problem confidence interval on that gain (0.0466-0.0695) excludes zero.