A novel method for accelerating LLM inference with reported speedups and benchmark comparisons, but presented as a preprint without peer review, limiting strength of evidence for clinical or production deployment claims.
Reported
Speedup magnitudeup to 3×
Model size (Uno)8B
Comparison model size (DiffusionG…26B