AUG 12, 2026 · PREPRINT
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
arXiv
This is an unrefereed preprint describing a machine learning framework for voice generation; it reports computational performance and comparative quality metrics but lacks clinical or medical evidence and has not undergone peer review.
Reported
Model parameters43.51 million
Minimum inference steps4 ODE steps