AUG 10, 2026 · PREPRINT
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training
arXiv
Early-stage machine learning method paper with experimental results on proprietary models; lacks peer review, independent validation, and causal evidence; results support a hypothesis rather than establish definitive superiority.
Reported
MATH500 accuracy (4B model, basel…0.630
MATH500 accuracy (4B model, WDL-O…0.685
MATH500 accuracy (1.7B model, bas…0.521