SEP 8, 2026 · PREPRINT
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
arXiv
A preprint describing a text-to-speech system with engineering design and benchmark comparisons, but lacking peer review, rigorous controlled evaluation, and clinical or diagnostic validation.
Reported
Time to first audioapproximately 200 ms
Non-streaming end-to-end RTF (sin…0.08
Non-streaming end-to-end RTF (eig…0.02