audio.cpp 0.4 Runs Higgs Audio v3 TTS Locally
audio.cpp 0.4 adds Higgs Audio v3 4B, Fish Audio S2 Pro, and Voxtral Realtime ASR to one C++/ggml engine, running flagship text-to-speech locally on CUDA with no Python.
audio.cpp 0.4 adds Higgs Audio v3 4B, Fish Audio S2 Pro, and Voxtral Realtime ASR to one C++/ggml engine, running flagship text-to-speech locally on CUDA with no Python.
audio.cpp 0.3 adds Supertonic 3, IndexTTS2, Irodori-TTS, and MOSS-TTS to one C++/ggml engine, running local text-to-speech and voice cloning with no Python and no cloud fees.
PocketTTS-RAVEN runs text-to-speech and voice cloning entirely in a browser tab, with no server, GPU, or account, at faster-than-realtime speed.
TTS Audio Suite shipped v5.0.0 for ComfyUI, adding Higgs Audio v3 voice cloning, Transformers 5, and Runtime Isolation for legacy engines.
A revamped open-source TTS benchmark now compares 46 text-to-speech models using objective scores and blind human voting, so creators can see which voices actually hold up.
Supertonic 3 is an open-weights, CPU-only TTS engine from Supertone with 31 languages, expression tags, and zero-shot voice cloning.
OpenReader v3.0 converts PDF, EPUB, DOCX, TXT, and Markdown files into synchronized read-along sessions or exported audiobooks, with multiple TTS providers and Docker deployment.
xAI launched the Grok Voice Agent API April 18 with standalone speech-to-text and text-to-speech endpoints priced at $0.10 per hour batch STT and $4.20 per million characters TTS.
The AI voice cloning market reaches $4.06 billion in 2026. We compare ElevenLabs, Voxtral TTS, and Fish Audio S2 on quality, pricing, latency, and self-hosting to help creators choose the right tool.
Mistral Voxtral TTS is a 4B parameter open-weights model that matches ElevenLabs quality in human evaluations, with 3-second voice cloning across 9 languages and self-hosting support.