Inflect-Micro-v2: Complete Text-to-Speech Synthesis in Just 9.36M Parameters
What happened
Researchers published Inflect-Micro-v2 on Hugging Face on July 25 -- a text-to-speech model handling the complete TTS pipeline including phoneme conversion, pitch duration, and acoustic decoding in just 9.36M parameters.
Context and impact
Most modern TTS models (ElevenLabs, Kokoro-82M, F5-TTS) have tens to hundreds of millions of parameters. Inflect-Micro-v2 demonstrates viable speech synthesis at a fraction of the compute requirements -- ideal for edge devices, wearables, and embedded systems.
Details
- 9.36M parameters -- compare: Kokoro-82M has ~82M, F5-TTS ~300M
- Complete pipeline: phoneme conversion + pitch duration + acoustic decoder in one model
- Released under permissive license on Hugging Face
- Hacker News score 148 -- strong community interest on July 25
Open original source
Hugging Face / Hacker News