Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is an unreviewed system development report describing a deterministic-prompt approach to single-speaker Greek text-to-speech synthesis using fine-tuned Parler-TTS and LoRA adaptation on limited audiobook data. The authors report acceptable word error rate (10.7%) and speaker consistency (MOS-C 4.24) but near-human quality (MOS-I 4.00) with no comparative arms or peer review.
Single-arm system development and technical validation study. Low-resource Greek language; single-speaker synthesis task. No human subjects enrolled.. Intervention: Deterministic prompting strategy combined with speaker-specific LoRA fine-tuning on 3.5 h audiobook data..
Word error rate (WER) of 10.7%, 2.9 points above ASR floor Mean opinion score for identity (MOS-I) of 4.00 versus 4.36 for human speech Speaker consistency (MOS-C) of 4.24 versus 4.30 for human reference
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A proof-of-concept system for low-resource Greek TTS with engineering validation (WER, MOS scores) but no comparative trial, no peer review, and no clinical application.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.