Life sciences · Preprint
arXiv · August 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
WDL-OPD is a proposed mixture-constrained co-training method for on-policy distillation that uses an anchor policy to generate rollouts and an auxiliary policy to evaluate them. In uncontrolled experiments on proprietary Qwen3 models, it improved MATH500 accuracy and code generation scores compared to single-policy baselines, but the authors themselves characterize the evidence as supporting a stabilization hypothesis rather than a universal causal claim, and several comparisons differ in curriculum or initialization.
Preprint. Qwen3 language models at 1.7B and 4B scales; no human subjects or clinical population.. Intervention: WDL-OPD mixture-constrained co-training method with anchor and auxiliary policies and geometric mixture matching to frozen teacher.. Compared with: Single-policy OPD configurations and related methods OPD² and W2S-OPD..
MATH500 accuracy raised from 0.630 to 0.685 at 4B scale and from 0.521 to 0.585 at 1.7B scale Code generation reached independently re-evaluated development scores of 0.637 and 0.375 Seven single-policy OPD configurations exhibited entropy growth or trajectory degradation; co-training avoided this in recorded experiments
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
Early-stage machine learning method paper with experimental results on proprietary models; lacks peer review, independent validation, and causal evidence; results support a hypothesis rather than establish definitive superiority.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The same feedback loop can nevertheless be unstable: each update changes both the policy and the states on which the next update is computed. We introduce WDL-OPD, a mixture-constrained co-training method with two trainable policies. An anchor policy generates every rollout, an auxiliary policy evaluates the same visited states, and a geometric mixture of their token distributions is matched to a frozen teacher by reverse KL. Both policies receive gradient. We show that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express. In recorded Qwen3 experiments at 1.7B and 4B scale, WDL-OPD produces the strongest student checkpoint in each of four scale-domain settings. It raises MATH500 accuracy from 0.630 to 0.685 at 4B and from 0.521 to 0.585 at 1.7B. In code generation, seven single-policy OPD configurations exhibit entropy growth or trajectory degradation, while co-training reaches independently re-evaluated development scores of 0.637 and 0.375. Because several comparisons differ in curriculum or initialization, these results support a stabilization hypothesis rather than a universal causal claim. We provide the exact training algorithm, failure evidence, and the controlled comparison matrix needed to test that hypothesis.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.