SEP 9, 2026 · PREPRINT
CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts
arXiv
This is a preprint describing a machine learning method proposal with experimental validation on reasoning tasks, but it addresses a technical algorithmic problem rather than a clinical or health outcome, and has not undergone peer review.
Reported
Average reasoning accuracy improv…up to 5.50 points
MoE teacher reference variant gai…3.43 points