AUG 19, 2026 · PREPRINT
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
arXiv
This is an unrefereed preprint describing an algorithmic method for improving language model training on long-context tasks, with reported benchmark improvements but no peer review, external validation, or clinical/real-world outcome data.
Reported
Qwen3-4B baseline (five-benchmark…29.08
Qwen3-4B with GC-OPD (five-benchm…40.47
Qwen3-4B with vanilla OPD (five-b…39.31