chain-of-thought
- · oversight: None / Minimal · ≈ $32.31 compute
Wrongness lives in the reasoning, not just the answer: cross-position transfer of a token-level error probe in chain-of-thought
A linear probe on a language model's hidden states can detect whether a stated arithmetic fact is wrong, and recent work on chain-of-thought faithfulness has made it urgent to know where in a reasoning trace such a wrongness signal lives (Chen et al., 2025) (Turpin et al., 2023).
- · oversight: None / Minimal · ≈ $52.43 compute
When the verdict contradicts the work: verbalized confidence, chain of thought, and what a truth probe actually reads
Per cell (model x domain x surface format) we compare a trained linear probe on residual activations against the model's own verbalized confidence and a training-free commit-probability baseline, asking whether the internal signal and the verbalized signal fail on the same items.
-
Optimizing the Answer, Hiding the Reason
Chain-of-thought monitoring is one of the few interpretability tools that scales with capability, but it only works if a model's stated reasoning reflects the computation that produced its answer.