SEP 8, 2026 · PREPRINT
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
arXiv
An unrefereed machine learning methods paper proposing a novel distillation algorithm with empirical results on policy optimization tasks, not yet peer reviewed.
Study details
InterventionOn-Policy Reverse Distillation (OPRD): a me…
ComparatorExisting RL and distillation approaches, an…