OpenAI GPT-Live: Full-Duplex ChatGPT Voice
OpenAI launched GPT-Live on July 8, 2026, a pair of full-duplex voice models that listen and speak at the same time, replacing Advanced Voice Mode for every ChatGPT user.
OpenAI launched GPT-Live on July 8, 2026, a pair of full-duplex voice models that listen and speak at the same time, replacing Advanced Voice Mode for every ChatGPT user.
Cohere released Transcribe Arabic, an open-weight Apache 2.0 speech-to-text model built for Arabic dialects, bilingual speech, and code-switching.
NVIDIA has released Audex, an open-weight audio-text model that does speech recognition, translation, text-to-speech, and general audio generation in one network.
PocketTTS-RAVEN runs text-to-speech and voice cloning entirely in a browser tab, with no server, GPU, or account, at faster-than-realtime speed.
Gladia shipped a command-line version of its speech-to-text platform, turning audio transcription into one terminal command with SRT and VTT subtitles, speaker diarization, and 100-plus languages.
ZeroLabs stitches six open-source models into one free browser studio for voice cloning, design, cleanup, sound effects, and transcription, running on Hugging Face ZeroGPU.
Interfaze's diffusion-gemma-asr-small is billed as the first open-source diffusion-based speech recognition model, refining random tokens into a transcript instead of decoding left to right.
audio.cpp packs more than a dozen audio AI models into one ggml-powered C++ runtime, running TTS, voice cloning, and music up to 5x faster than Python.
DeepL has acquired Mixhalo, the real-time audio platform built for concerts and conferences, pushing AI voice translation into live events.
Cartesia launched Sonic-3.5 and Ink-2 on June 15, 2026, claiming the top streaming voice AI on both speaking and listening, built for real-time voice agents.
TTS Audio Suite shipped v5.0.0 for ComfyUI, adding Higgs Audio v3 voice cloning, Transformers 5, and Runtime Isolation for legacy engines.
The best AI music generators in 2026, from Suno and Udio to ElevenLabs Music, Google Lyria, and Stable Audio. We compare standout features, pricing, and commercial-use rights so you can pick the right tool for songs, scores, or royalty-free background music.
A new open-source Audio-Reactive LoRA for LTX 2.3 turns a still image and an audio track into a music-driven clip, with motion synced to the beat.
Suno's new Advanced Split mode lets Premier subscribers extract any of nearly 100 instruments from a track, from a full drum kit to vocals.
A revamped open-source TTS benchmark now compares 46 text-to-speech models using objective scores and blind human voting, so creators can see which voices actually hold up.
A new research benchmark published June 1, 2026 reveals that state-of-the-art AI music detectors systematically fail on hybrid productions, the kind created by most real-world music producers using tools like Suno or Udio.
OpenMOSS published the MOSS-Audio technical report on June 1, 2026, documenting four open-source audio-language models that achieve benchmark scores rivaling systems three to four times their size.
A developer used Stability AI Stable Audio 3 Medium to generate 15,834 free audio samples: 10,359 drum one-shots and 5,475 pitched instrument recordings, available for immediate download.
Baidu open-sourced NAVA, a 6.3B parameter joint audio-video model that generates 720p video with synced dual-channel audio in a single pass.
parakeet.cpp ports NVIDIA Parakeet automatic speech recognition models to ggml, eliminating the Python runtime entirely — with byte-identical output to NeMo at up to 1.86x faster throughput.
A team of 10 researchers published Dasheng AudioGen on May 27, 2026, a unified model that generates complete audio scenes from text descriptions.
ElevenLabs released Music v2 on May 26, 2026 with dynamic genre transitions and up to 50% lower prices. Live now on ElevenMusic and ElevenCreative.
Orchestria launches May 25 with a multi-agent AI music engine that exposes drums, bass, and melody as editable stems. Free, Pro $8, Maestro $25.
Developer Shukant Pal built Pretzel at the Google I/O hackathon on May 24, 2026: a live web-based music sequencer where an AI agent controls what everyone hears. No signup required.