deception-honesty
-
Weak-label finetuning does not overwrite a strong model's clean-label direction, even as its output collapses
When a strong language model is finetuned on labels from an unreliable supervisor and then starts giving wrong answers, it is natural to assume the bad labels have corrupted what the model internally knows. We test that assumption directly and find it false in our setting.