Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
TontaubeV1 is a streaming text-to-speech model designed to balance prosody quality with low-latency GPU inference. Performance is claimed via an LLM-as-a-judge benchmark on audiobook reading; the work has not undergone peer review and lacks human perceptual validation or rigorous comparative methodology.
Preprint. Text-to-speech generation for English, German, and multilingual support; audiobook reading as benchmark; accepts up to one minute reference audio for voice conditioning.. Intervention: TontaubeV1 text-to-speech model using hierarchical DualCodec, semantic stream prediction via Qwen3-1.7B transformer, and three acoustic refinement transformers derived from Qwen3-0.6B.. Compared with: ElevenLabs Flash v2.5, Fish Audio S2 Pro, Gradium API (April 2026), and Cartesia Sonic 3; evaluation on LLM-as-a-judge audiobook-reading benchmark..
Time to first audio on single RTX 5090: approximately 200 ms Non-streaming end-to-end real-time factor (RTF): 0.08 for one input, 0.02 across eight concurrent inputs On LLM-as-a-judge audiobook benchmark, TontaubeV1 matches ElevenLabs Flash v2.5 and outperforms Fish Audio S2 Pro, Gradium API (April 2026), and Cartesia Sonic 3 on prosody
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A preprint describing a text-to-speech system with engineering design and benchmark comparisons, but lacking peer review, rigorous controlled evaluation, and clinical or diagnostic validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptual quality typically comes at increased computational cost and latency. We present TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU. Speech is encoded by the hierarchical DualCodec representation at 12.5 Hz, which separates a semantic stream from successive acoustic refinements. Our design assumes that prosodic structure is largely established when the semantic stream is generated, and allocates capacity accordingly: a Qwen3-1.7B-derived transformer predicts that stream and thereby the utterance duration, while three progressively smaller Qwen3-0.6B-derived transformers each add one acoustic refinement. Text is tokenized per character rather than by subword. Paired text and audio markers at shared positions support long-form generation with bounded context, and overlapping DualCodec reconstructions are mapped into the VibeVoice acoustic latent space and decoded causally, enabling streaming despite DualCodec's noncausal decoder. The model accepts up to one minute of reference audio for voice conditioning and is designed primarily for English and German, with additional multilingual support. The four predictors total 2.9B parameters; on a single RTX 5090 the streaming path reaches approximately 200 ms to first audio. In separate non-streaming measurements, the end-to-end real-time factor (RTF) is 0.08 for one input and the aggregate RTF is 0.02 across eight concurrent inputs. On our LLM-as-a-judge audiobook-reading benchmark, TontaubeV1 matches ElevenLabs Flash v2.5 and outperforms Fish Audio S2 Pro, the April 2026 Gradium API, and Cartesia Sonic 3 on prosody. The model weights are released on Hugging Face under the Tontaube Community Model License 1.0.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.