MIT flow-matching TTS with expressive multilingual voices at consumer-GPU speed.
Catalog snapshot
Fine as CPU inference, but a 6 GB GPU makes it pleasant.
Fetched 1
F5-TTS is an MIT flow-matching voice model with natural prosody, fast inference, and strong multilingual coverage. The HuggingFace Space made it the default open narration engine. Runs on 6 GB VRAM or patient CPU.
flow-matching TTS · speed control · 28-language support · voice conversion