Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint proposes Negative Self-Distillation (NSD), a novel training framework for large language models that aims to improve reasoning by learning to avoid flawed reasoning patterns rather than imitating correct solutions. The authors report empirical outperformance of NSD over On-Policy Self-Distillation (OPSD) and reinforcement learning baselines, but the work remains unrefereed and lacks independent validation or specification of effect sizes.
Preprint. Large language models. Intervention: Negative Self-Distillation (NSD): training framework that optimizes LLMs by diverging from self-generated flawed reasoning using a dynamic gating mechanism to isolate reasoning-critical tokens. Compared with: On-Policy Self-Distillation (OPSD) and label-free, self-bootstrapping reinforcement learning baselines.
NSD consistently outperforms OPSD and label-free, self-bootstrapping RL baselines on complex reasoning tasks OPSD degrades LLM performance on complex reasoning by suppressing uncertainty and penalizing self-corrective behaviors NSD uses a dynamic gating mechanism to isolate reasoning-critical tokens and avoid degrading foundational language capabilities
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
Not applicable; this is a machine learning methods paper without direct clinical or healthcare application described in the source.
This is an unrefereed preprint describing a novel machine learning method with empirical comparisons to baselines, but lacks peer review, independent validation, and clinical or real-world deployment evidence.
As stated by the source record.
Not applicable; this is a machine learning methods paper without direct clinical or healthcare application described in the source.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.