This preprint reports an empirical finding from language model experiments showing sparse supervision can match dense supervision, but lacks peer review, uses surrogate endpoints (reasoning task performance), and presents a phenomenon requiring mechanistic explanation rather than answering a clinical or definitive practical question.
Reported
Supervision sparsity0.05%
Tokens per trajectory1 to 2
Teacher-student configurations te…9