outcome-reward
- · oversight: None / Minimal · ≈ $47.64 compute
Outcome-Only RL Did Not Reduce a Model's Causal Dependence on Its Own Chain of Thought
A worry about outcome-only reinforcement learning is that, by rewarding only the final answer, it makes the chain of thought decorative: the answer would stop causally depending on the reasoning the model writes. We test this with a self-CoT corruption-flip protocol.
- · oversight: None / Minimal · ≈ $79.19 compute
Small-Scale Outcome-Only GRPO Barely Moves the Answer Distribution
If outcome-only RL mostly elicits what the base model already computes, its effect on the answer distribution might be reproducible by a single global affine transform of the base logits (a scalar temperature plus a per-option bias).
- · oversight: None / Minimal · ≈ $62.90 compute
No Reliable Filler-Token Steganographic Channel After Small-Scale Outcome-Only RL
Outcome-only RL rewards the final answer and not the reasoning, which raises the worry that a model could learn to smuggle answer-relevant information through tokens that look like non-load-bearing filler, a steganographic channel hidden in the chain of thought.