Sonar Brings Audio Search to AI Agents and Podcast Creators
Sonar entered public beta offering natural-language search across podcasts, news, earnings calls, and radio archives for AI agents. Free tier: 500 queries/month.
Sonar entered public beta offering natural-language search across podcasts, news, earnings calls, and radio archives for AI agents. Free tier: 500 queries/month.
Spotify's new Studio desktop app uses an AI agent to generate private, personalized podcasts from your calendar, email, and web browsing. In research preview across 20+ markets.
Spotify announced a new ElevenLabs-powered audiobook creation tool at its May 21 Investor Day, with new audiobook subscription plans expected later in 2026.
Wrap an existing chat agent (Claude, GPT, Gemini) with ElevenLabs Speech Engine voice in roughly 30 minutes. Step-by-step tutorial covers Node and Python SDKs, WebRTC token flow, turn detection, interruption, and the upgrade path to ElevenAgents.
Researchers at Tampere University published a preprint introducing ACAD, an AI system that adapts what it treats as noise based on the acoustic scene it detects.
Stability AI released Stable Audio 3 on May 20, 2026, a family of open-weights latent diffusion models that generate up to 6 minutes and 20 seconds of audio from a text prompt, including on a MacBook Pro M4.
StemDeck v0.5.0 splits any track into vocals, drums, bass, guitar, piano, and other elements. Runs locally, free, no account required.
ComfyUI-DramaBox added LoRA weight injection on May 16, letting creators load custom voice personalities into workflows without model reloads.
OpenAI quietly bought voice-cloning startup Weights.gg, dispersed its six-person team across existing groups, and folded the tech into Voice Engine work.
Supertonic 3 is an open-weights, CPU-only TTS engine from Supertone with 31 languages, expression tags, and zero-shot voice cloning.
Break-the-Beat! is a new AI model that renders drum MIDI patterns as realistic audio using a reference recording for timbre. Researchers released a demo with dozens of examples spanning Speed Metal, Funk Rock, and electronic drum kits.
ScenemaAI released Scenema Audio on Hugging Face and GitHub, an open-weights expressive TTS and zero-shot voice cloning model built on the audio half of Lightricks LTX-2. MIT inference code, 13 languages, real-time on a 24 GB GPU.
Google partners with Believe and TuneCore to route Flow Music, its Lyria 3 Pro song studio, to indie artists. Ambassador program meets Google's product team weekly.
ElevenLabs crossed $500 million in ARR on May 5, 2026, announcing BlackRock, Nvidia, Jamie Foxx, and Eva Longoria as new investors. The $11B valuation makes it the most-funded independent voice AI company.
Ableton Live 12.4 shipped May 5, 2026 as a free update for Live 12. The standout feature is Link Audio, which streams audio wirelessly between two devices on your local network without extra hardware.
Jamie Pine's Voicebox brings seven TTS engines and voice cloning to your local machine. Free, open source, and entirely offline.
RODE announced RODECaster Studio at NAB 2026 on April 17, 2026 -- a new Mac and Windows desktop app that brings AI-powered editing to podcast post-production.
Google launched a native Gemini app for Mac on April 16, 2026, bringing image, video, and music generation directly to the desktop for the first time.
Researchers published Darwin-TTS on April 15, 2026, a text-to-speech model that adds emotional expression to AI voice without any training, fine-tuning, or new data.
Google released Gemini 3.1 Flash TTS on April 15, 2026, a text-to-speech model that outperforms ElevenLabs v3 in quality benchmarks while offering a generous free tier.
Splice launched Variations and Craft on April 15, 2026, letting producers remix any sample from its 3-million-sound library while automatically compensating the original creator.
ComfyUI now supports Sonilo via Partner Nodes, letting creators generate full-length soundtracks that sync to video footage frame by frame.
llama.cpp release b8769 adds audio multimodal support for Qwen3-Omni and Qwen3-ASR models, bringing local speech recognition and audio understanding to consumer hardware.
OpenBMB released VoxCPM2, a 2 billion parameter text-to-speech model that runs on 8GB VRAM, supports 30 languages at 48kHz, and can design voices from natural language descriptions.