Life sciences · Preprint
arXiv · September 4, 2026
Raises a question worth testing. It does not answer one.
This preprint reports a computational study comparing four optimizers (AdamW, ADANA, Muon, SOAP) across overtraining factors from 1× to 256× on language models of 51M to 253M parameters. The authors find that optimal hyperparameters and relative optimizer performance vary substantially with training horizon, and that ADANA's advantage over AdamW persists and grows with training length, approaching predictions from DANA theory. However, the work is unreviewed, limited to a narrow model scale and undescribed datasets, and does not establish clinical or production-system validity.
Computational comparison study. Language models of specified parameter ranges trained under controlled overtraining conditions.. Intervention: ADANA, Muon, and SOAP optimizers with tuned hyperparameters (log-time weight decay, momentum cooldown, fixed memory schedules).. Compared with: AdamW optimizer with fixed and separately tuned fixed memory across horizons..
Optimal learning rate schedule can reverse across the overtraining axis Best weight decay coefficient scales approximately as sqrt(OT) Longer training horizons generally favor longer fixed memory in optimizers
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
Exploratory computational study comparing optimizer scaling behaviour across training horizons; no clinical or established-field endpoints, no peer review, and findings are algorithmic rather than validated against independent data or clinical outcomes.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
We investigate how optimizers scale across the overtraining axis and show that relative optimizer performance and optimal hyperparameters change substantially with training horizon. In particular, we study how matrix-preconditioned methods (Muon and SOAP) and a momentum-scheduled method (ADANA) scale relative to AdamW. We compare these four optimizers across models from 51M to 253M parameters and overtraining (OT) factors from 1x to 256x, sweeping the base learning rate at every setting. The preferred learning rate schedule can reverse across the overtraining axis, the best weight decay coefficient scales approximately as sqrt(OT), and longer horizons generally favor longer fixed memory. ADANA's scaling advantage over AdamW persists after tuning AdamW's fixed memory separately at each horizon. Log-time weight decay and momentum cooldown provide substantial gains for ADANA that compound as training increases. With this treatment, ADANA outscales AdamW with an exponent advantage close to that predicted by DANA theory on power-law random features. Muon and SOAP instead provide roughly constant token-efficiency advantages over AdamW across most of the measured range, although SOAP may gain further at the highest overtraining factors. ADANA begins behind both matrix-preconditioned optimizers but closes the gaps as training increases, surpassing Muon and becoming competitive with SOAP at our highest OT factors. These results establish training horizon as an essential axis for optimizer evaluation and design.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.