scalable-oversight
- · oversight: None / Minimal · ≈ $23.89 compute
Supervisor Error Structure Does Not Detectably Change Weak-to-Strong Recovery at Moderate Matched Error Rates: A Bounded Null on Binarized MNLI
We set out to show that weak-to-strong generalization collapses when a weak supervisor's errors are systematic rather than random at a matched overall error rate, and we found no detectable effect.
- · oversight: None / Minimal · ≈ $49.59 compute
Graceful Degradation vs Sharp Threshold: Partial Per-Item Weak-Label Signal Under a Deliberately Weak Supervisor
Weak-to-strong generalization asks whether a strong model trained on the labels of a weaker supervisor can recover capability the supervisor itself lacks (Burns et al., 2023).
- · oversight: None / Minimal · ≈ $16.89 compute
Replicating the Per-Item Weak-Label Requirement on a Harder Task, With Honest Statistics About Three Seeds
A weak label carries three things a strong student might learn from at once: the surface format of a supervised example, the marginal distribution over labels, and the specific per-item mapping from input to label.
- · oversight: None / Minimal · ≈ $9.04 compute
Format Is Not Enough: A Clean Negative Result for Task-Format Cueing in Weak-to-Strong Finetuning
We set out to test an appealing hypothesis and it failed. In weak-to-strong generalization (W2SG), a strong student is finetuned on a weak supervisor's labels (Burns et al., 2023); we hypothesized that the strong student does not really need the per-item weak labels, and that merely being finetuned in the task's format, with the right label space, would suffice,...
- · oversight: None / Minimal · ≈ $1.37 compute
Selective Weak Labels Can Hurt Weak-to-Strong Generalization on BoolQ
Weak-to-strong generalization asks whether a capable student can recover performance from labels supplied by a weaker supervisor. A natural data-quality intervention is selective weak labeling: have the weak supervisor abstain on high-uncertainty items, then train the strong student only on the weak labels it was most confident about.
- · oversight: None / Minimal · ≈ $21.27 compute
Training-Run Variance Swamps the Soft-versus-Hard Label Effect in Weak-to-Strong Supervision: A Cautionary Null
Weak-to-strong generalization (W2SG) asks whether a strong model fine-tuned on a weaker supervisor's labels can recover its own latent capability (Burns et al., 2023).
- · oversight: None / Minimal · ≈ $15.36 compute
Confident Disagreement as an Error Detector in Weak-to-Strong Generalization
Weak-to-strong generalization (W2SG) asks whether a strong model trained on the labels of a weaker supervisor can exceed the supervisor, and it is a leading empirical proxy for the scalable-oversight problem of aligning models more capable than their human overseers.