Deep Dive
Jul 28, 2026
VoiceHop 2.0 brings sub-second, voice-preserving AI translation to any video, stream, or call as a browser extension. Here is what it does and how real-time translation compares to async dubbing tools.
Deep Dive
Jul 28, 2026
Fish Audio raised $52M and made S2.1 Pro free through August 31: production-grade voice cloning across 83 languages at roughly 90ms latency, plus an open-source path via Fish Speech.
Audio
Jul 24, 2026
audio.cpp 0.4 adds Higgs Audio v3 4B, Fish Audio S2 Pro, and Voxtral Realtime ASR to one C++/ggml engine, running flagship text-to-speech locally on CUDA with no Python.
Deep Dive
Jul 22, 2026
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a hosted text-to-speech model that topped the Artificial Analysis TTS leaderboard in July 2026.
Deep Dive
Jul 14, 2026
audio.cpp 0.3 adds Supertonic 3, IndexTTS2, Irodori-TTS, and MOSS-TTS to one C++/ggml engine, running local text-to-speech and voice cloning with no Python and no cloud fees.
Deep Dive
Jul 13, 2026
Kyutai and Mirelo released MuScriptor, an open-weight model that transcribes a full song mix into separate, editable MIDI tracks, one per instrument.
Deep Dive
Jul 9, 2026
aria is a dependency-free native runtime that runs the full Stable Audio 3 text-to-music pipeline on ordinary GPUs, CPU-only laptops, and an 8GB Raspberry Pi 5, no Python required.
AI
Jul 8, 2026
Willow launched two speech-to-text models: Frontier Mini, a free unlimited tier, and Frontier Pro, a faster paid model built for power users.
Deep Dive
Jul 8, 2026
OpenAI launched GPT-Live on July 8, 2026, a pair of full-duplex voice models that listen and speak at the same time, replacing Advanced Voice Mode for every ChatGPT user.
AI
Jul 7, 2026
Cohere released Transcribe Arabic, an open-weight Apache 2.0 speech-to-text model built for Arabic dialects, bilingual speech, and code-switching.
AI
Jul 7, 2026
NVIDIA has released Audex, an open-weight audio-text model that does speech recognition, translation, text-to-speech, and general audio generation in one network.
Audio
Jul 7, 2026
PocketTTS-RAVEN runs text-to-speech and voice cloning entirely in a browser tab, with no server, GPU, or account, at faster-than-realtime speed.
Deep Dive
Jul 6, 2026
Gladia shipped a command-line version of its speech-to-text platform, turning audio transcription into one terminal command with SRT and VTT subtitles, speaker diarization, and 100-plus languages.
Deep Dive
Jul 2, 2026
ZeroLabs stitches six open-source models into one free browser studio for voice cloning, design, cleanup, sound effects, and transcription, running on Hugging Face ZeroGPU.
Deep Dive
Jul 2, 2026
Interfaze's diffusion-gemma-asr-small is billed as the first open-source diffusion-based speech recognition model, refining random tokens into a transcript instead of decoding left to right.
News
Jun 27, 2026
audio.cpp packs more than a dozen audio AI models into one ggml-powered C++ runtime, running TTS, voice cloning, and music up to 5x faster than Python.
News
Jun 17, 2026
DeepL has acquired Mixhalo, the real-time audio platform built for concerts and conferences, pushing AI voice translation into live events.
Deep Dive
Jun 15, 2026
Cartesia launched Sonic-3.5 and Ink-2 on June 15, 2026, claiming the top streaming voice AI on both speaking and listening, built for real-time voice agents.
AI
Jun 14, 2026
TTS Audio Suite shipped v5.0.0 for ComfyUI, adding Higgs Audio v3 voice cloning, Transformers 5, and Runtime Isolation for legacy engines.
Deep Dive
Jun 13, 2026
The best AI music generators in 2026, from Suno and Udio to ElevenLabs Music, Google Lyria, and Stable Audio. We compare standout features, pricing, and commercial-use rights so you can pick the right tool for songs, scores, or royalty-free background music.
Open Source
Jun 12, 2026
A new open-source Audio-Reactive LoRA for LTX 2.3 turns a still image and an audio track into a music-driven clip, with motion synced to the beat.
AI
Jun 11, 2026
Suno's new Advanced Split mode lets Premier subscribers extract any of nearly 100 instruments from a track, from a full drum kit to vocals.
Audio
Jun 8, 2026
A revamped open-source TTS benchmark now compares 46 text-to-speech models using objective scores and blind human voting, so creators can see which voices actually hold up.
Deep Dive
May 31, 2026
OpenMOSS published the MOSS-Audio technical report on June 1, 2026, documenting four open-source audio-language models that achieve benchmark scores rivaling systems three to four times their size.