AUG 10, 2026 · PREPRINT
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
arXiv
A novel systems technique for optimizing inference in looped language models, demonstrated on two specific models with throughput and latency metrics, but lacking peer review, clinical validation, or comparison to established baselines in a controlled setting.
Reported
Theoretical speed-up realizedup to 99%
Offline throughput improvement1.5–1.9×
Normalized latency reduction45–90%