Large-scale computational study proposing novel hyperparameter scaling laws for sparse models, but lacks peer review, clinical/patient outcomes, and independent validation on held-out test cases beyond one ultra-sparse configuration.
Reported
Pre-training runs1,800
Maximum model size (total non-emb…6B
Total tokens processedapproximately 20 trillion