Exploratory computational study comparing optimizer scaling behaviour across training horizons; no clinical or established-field endpoints, no peer review, and findings are algorithmic rather than validated against independent data or clinical outcomes.
Reported
Model parameter range51M to 253M
Overtraining factor range1x to 256x
Weight decay scaling relationshipsqrt(OT)