Swiftlet: Run an 80B Qwen Locally on a Mac or iPhone
Swiftlet runs the real Qwen3-Next-80B on a Mac in 4.3 GB of RAM and a 35B model on an iPhone, using expert streaming to keep only active experts in memory.
In-depth analysis and deep dives into the AI tools shaping creative work.
Swiftlet runs the real Qwen3-Next-80B on a Mac in 4.3 GB of RAM and a 35B model on an iPhone, using expert streaming to keep only active experts in memory.
Google shipped model routing for Cloud API Gateway on August 4, 2026: one OpenAI-compatible endpoint that forwards each request to Gemini, Claude, or an OpenAI GPT model.
Velokey is a new pay-as-you-go API gateway that unifies Kling, Veo, Sora, Seedance, plus image and LLM models behind one OpenAI-compatible endpoint.
NVIDIA released open weights for VoiceChat-11B, the first open full-duplex voice model that calls tools mid-conversation. Here is how it works and how to self-host it.
If you code under an NDA or on a private repo, an on-device voice coding tool is the only safe way to talk to your AI agent. SKI, Spokenly, and VoiceMode keep every word local. Here is how the leading voice coding tools compare.
FFmpeg 9.0 Lei shipped August 3, 2026, the most GPU-focused release in years: animated WebP decode, Vulkan 360 filtering, ProRes RAW on Apple, and AMD AMF HDR filters.
Cloudflare shipped three inference optimizations that make Kimi K2.6 and GLM 5.2 roughly 30% cheaper to run on Workers AI, and its benchmark tables show the answers do not change.
Alibaba announced Qwen3.8-Max on August 3, 2026: a 2.4 trillion-parameter coding flagship with a 1 million-token context and a promised open-weights release.
As of August 2, 2026, the EU AI Act Article 50 transparency rules are enforceable. Here is what creators must disclose and how to label AI-generated content.
LoRA Dataset Studio is a free, self-hosted web app that runs the entire LoRA training lifecycle inside a single browser tab, from one reference photo to a trained, ranked model.
xAI expanded Grok Imagine Video 1.5 with native 1080p, text-to-video, and up to seven locked image and voice references, live August 1 on web, iOS, Android, and the API.
audio.cpp 0.5 turns a pure C++ engine into a full local audio studio: expressive DramaBox TTS, Confucius4 cross-lingual voice transfer, music, and AMD ROCm support, all with no Python.
Ideogram P-Image starts at $0.003 per image with Ideogram signature text rendering. Here is how it compares to GPT-Image, FLUX, Ideogram 4.0, and Qwen on price, speed, and in-image text.
Y Combinator open-sourced QM, an MIT-licensed multiplayer agent harness that gives every teammate an isolated workspace and drives Claude Code, Codex, OpenCode, and Pi from one core.
DeepSeek shipped the official production build of V4 Flash on July 31, 2026: an open-weight, MIT-licensed 284B-parameter MoE model with a 1M token context, native OpenAI Responses API and Codex support, and API output at $0.28 per million tokens.
MiniMax released H3 on July 31, 2026: native 2K video with built-in stereo sound, all-modal input, and a price it says runs at a third of rival models.
Figma Make just added a properties panel and annotations, letting designers edit AI-generated apps visually instead of describing every tweak in a chat box.
NVIDIA open-sourced NOOA, a model-agnostic framework that turns an AI agent into an ordinary Python object. Here is how it works, the benchmarks, and how to build your first agent.
A step-by-step workflow to design a logo and starter brand identity with AI in about an hour, using a text-capable generator, a vectorizer, and free type and color tools.
Perplexity Projects turns Computer into a shared, multiplayer workspace for people and AI agents, with a persistent file system and Brain memory that carries context between tasks.
Cursor shipped a dedicated iPad app on July 29, 2026, letting you run, review, and merge AI coding agents with split screen, pinned sidebar chats, and Apple Pencil markup, no laptop required.
Google DeepMind released Lyria 3.5 in Flow Music on July 29, 2026, upgrading vocals, lyrics, musicality, and tempo control. Here is what changed and how it compares to Suno and Lyria 3 Pro.
Replit launched Replit Design on July 29, 2026, an AI creative suite that takes an idea to a live app with Mobbin references, brand systems, and no handoff. How it works and how it compares.
TurboFieldfare is an open-source Swift and Metal engine that runs Google Gemma 4 26B in about 2GB of RAM on any M-series Mac by streaming experts from SSD.