Deep Dive
May 31, 2026
OpenMOSS published the MOSS-Audio technical report on June 1, 2026, documenting four open-source audio-language models that achieve benchmark scores rivaling systems three to four times their size.
Audio
May 30, 2026
A developer used Stability AI Stable Audio 3 Medium to generate 15,834 free audio samples: 10,359 drum one-shots and 5,475 pitched instrument recordings, available for immediate download.
News
May 29, 2026
Baidu open-sourced NAVA, a 6.3B parameter joint audio-video model that generates 720p video with synced dual-channel audio in a single pass.
Deep Dive
May 28, 2026
parakeet.cpp ports NVIDIA Parakeet automatic speech recognition models to ggml, eliminating the Python runtime entirely — with byte-identical output to NeMo at up to 1.86x faster throughput.
Deep Dive
May 27, 2026
A team of 10 researchers published Dasheng AudioGen on May 27, 2026, a unified model that generates complete audio scenes from text descriptions.
ai-music
May 26, 2026
ElevenLabs released Music v2 on May 26, 2026 with dynamic genre transitions and up to 50% lower prices. Live now on ElevenMusic and ElevenCreative.
News
May 25, 2026
Orchestria launches May 25 with a multi-agent AI music engine that exposes drums, bass, and melody as editable stems. Free, Pro $8, Maestro $25.
Audio
May 24, 2026
Developer Shukant Pal built Pretzel at the Google I/O hackathon on May 24, 2026: a live web-based music sequencer where an AI agent controls what everyone hears. No signup required.
News
May 21, 2026
Sonar entered public beta offering natural-language search across podcasts, news, earnings calls, and radio archives for AI agents. Free tier: 500 queries/month.
Audio
May 21, 2026
Spotify announced a new ElevenLabs-powered audiobook creation tool at its May 21 Investor Day, with new audiobook subscription plans expected later in 2026.
Audio
May 21, 2026
Spotify's new Studio desktop app uses an AI agent to generate private, personalized podcasts from your calendar, email, and web browsing. In research preview across 20+ markets.
Deep Dive
May 21, 2026
Wrap an existing chat agent (Claude, GPT, Gemini) with ElevenLabs Speech Engine voice in roughly 30 minutes. Step-by-step tutorial covers Node and Python SDKs, WebRTC token flow, turn detection, interruption, and the upgrade path to ElevenAgents.
Audio
May 21, 2026
Researchers at Tampere University published a preprint introducing ACAD, an AI system that adapts what it treats as noise based on the acoustic scene it detects.
Deep Dive
May 20, 2026
Stability AI released Stable Audio 3 on May 20, 2026, a family of open-weights latent diffusion models that generate up to 6 minutes and 20 seconds of audio from a text prompt, including on a MacBook Pro M4.
Deep Dive
May 18, 2026
StemDeck v0.5.0 splits any track into vocals, drums, bass, guitar, piano, and other elements. Runs locally, free, no account required.
Deep Dive
May 15, 2026
ComfyUI-DramaBox added LoRA weight injection on May 16, letting creators load custom voice personalities into workflows without model reloads.
OpenAI
May 15, 2026
OpenAI quietly bought voice-cloning startup Weights.gg, dispersed its six-person team across existing groups, and folded the tech into Voice Engine work.
tts
May 15, 2026
Supertonic 3 is an open-weights, CPU-only TTS engine from Supertone with 31 languages, expression tags, and zero-shot voice cloning.
Deep Dive
May 14, 2026
Break-the-Beat! is a new AI model that renders drum MIDI patterns as realistic audio using a reference recording for timbre. Researchers released a demo with dozens of examples spanning Speed Metal, Funk Rock, and electronic drum kits.
Open Source
May 13, 2026
ScenemaAI released Scenema Audio on Hugging Face and GitHub, an open-weights expressive TTS and zero-shot voice cloning model built on the audio half of Lightricks LTX-2. MIT inference code, 13 languages, real-time on a 24 GB GPU.
News
May 6, 2026
Google partners with Believe and TuneCore to route Flow Music, its Lyria 3 Pro song studio, to indie artists. Ambassador program meets Google's product team weekly.
Deep Dive
May 5, 2026
ElevenLabs crossed $500 million in ARR on May 5, 2026, announcing BlackRock, Nvidia, Jamie Foxx, and Eva Longoria as new investors. The $11B valuation makes it the most-funded independent voice AI company.
Deep Dive
May 5, 2026
Ableton Live 12.4 shipped May 5, 2026 as a free update for Live 12. The standout feature is Link Audio, which streams audio wirelessly between two devices on your local network without extra hardware.
Audio
Apr 20, 2026
Jamie Pine's Voicebox brings seven TTS engines and voice cloning to your local machine. Free, open source, and entirely offline.