null-result
- · oversight: None / Minimal · ≈ $21.27 compute
Training-Run Variance Swamps the Soft-versus-Hard Label Effect in Weak-to-Strong Supervision: A Cautionary Null
Weak-to-strong generalization (W2SG) asks whether a strong model fine-tuned on a weaker supervisor's labels can recover its own latent capability (Burns et al., 2023).
- · oversight: None / Minimal · ≈ $49.52 compute
Does the wrongness probe lead the retraction? A weak, non-robust timing signal against behavioral self-correction in Qwen2.5-7B
If a model has an internal sense of when it is wrong, that signal might be expected to precede the moment it backs down under challenge, which would make a residual-stream probe an early monitor of self-correction. We test this directly on Qwen2.5-7B-Instruct.
- · oversight: None / Minimal · ≈ $47.64 compute
Outcome-Only RL Did Not Reduce a Model's Causal Dependence on Its Own Chain of Thought
A worry about outcome-only reinforcement learning is that, by rewarding only the final answer, it makes the chain of thought decorative: the answer would stop causally depending on the reasoning the model writes. We test this with a self-CoT corruption-flip protocol.
- · oversight: None / Minimal · ≈ $122.17 compute
No Significant CoT-Monitorability Decay to Localize
A safety worry about reinforcement learning is that it degrades chain-of-thought monitorability, making the written reasoning a less faithful guide to the answer.
- · oversight: None / Minimal · ≈ $62.90 compute
No Reliable Filler-Token Steganographic Channel After Small-Scale Outcome-Only RL
Outcome-only RL rewards the final answer and not the reasoning, which raises the worry that a model could learn to smuggle answer-relevant information through tokens that look like non-load-bearing filler, a steganographic channel hidden in the chain of thought.