DeepSeek-V4-Flash Makes LLM Steering Practical
DeepSeek-V4-Flash is the first local model competitive with frontier AI, making LLM activation steering practical for the first time. A guide for creators.
DeepSeek-V4-Flash is the first local model competitive with frontier AI, making LLM activation steering practical for the first time. A guide for creators.
NVIDIA's SANA-WM is a 2.6 billion-parameter open-source world model that generates 720p, 60-second video with 6-degree-of-freedom camera control on a single GPU.
ComfyUI-DramaBox added LoRA weight injection on May 16, letting creators load custom voice personalities into workflows without model reloads.
Supertonic 3 is an open-weights, CPU-only TTS engine from Supertone with 31 languages, expression tags, and zero-shot voice cloning.
Black Forest Labs' FLUX.2-klein-9B now integrates with ComfyUI via NKD Klein Tools v1.7.0, enabling sub-second image generation with just 4 inference steps on an RTX 4090.
holaOS 0.1 ships Dashboard, Sub Agents, and Multi Workspaces for managing parallel AI workflows on macOS.
Anthropic's Claude Code ported Bun from Zig to Rust in nine days: 1,009,257 lines, 99.8 percent test pass, 13,000-plus unsafe blocks. The case study.
IBM has released Granite Embedding Multilingual R2: a pair of Apache 2.0 embedding models with 32K context, 200+ languages, and a top MTEB score under 100M parameters. A drop-in swap for paid commercial embedding APIs.
Three significant open-source models arrived in ComfyUI on May 14, 2026: VOID for video object deletion, BiRefNet for image segmentation, and Gemma 4 for multimodal reasoning.
Open-source Claude Code skill turns one image into a fully meshed 3D scene with separate object meshes, Gaussian splat background, and ambient plus physics audio in under five minutes.
LTX Director v1.3.0 dropped today: a single ComfyUI node that replaces your entire LTX 2.3 toolkit. Timeline editing, prompt relay, 8 keyframes, and custom audio, all in one place.
OpenReader v3.0 converts PDF, EPUB, DOCX, TXT, and Markdown files into synchronized read-along sessions or exported audiobooks, with multiple TTS providers and Docker deployment.
ScenemaAI released Scenema Audio on Hugging Face and GitHub, an open-weights expressive TTS and zero-shot voice cloning model built on the audio half of Lightricks LTX-2. MIT inference code, 13 languages, real-time on a 24 GB GPU.
Cline released @cline/sdk on May 13, 2026, an open-source TypeScript runtime that lets developers embed the same coding agent powering Cline.
Nous Research's Hermes Agent reached 140,000 GitHub stars in three months. NVIDIA featured it for RTX creators who want a local AI agent that writes its own skills.
OpenSquilla v0.1.0 launched May 12, 2026: an Apache 2.0 microkernel runtime that smart-routes agent traffic across model tiers for 60 to 80 percent token savings.
TencentARC Pixal3D generates 3D models that stay pixel-aligned with your input image. Accepted to SIGGRAPH 2026, with a live HuggingFace demo and open-source code.
OpenBMB shipped MiniCPM-V 4.6, a 1.3B Apache 2.0 vision-language model that runs on iPhone, Android, and HarmonyOS and leads the sub-2B open-weights pack.
HiDream-O1-Image is an open-source 8B pixel transformer that beats FLUX.2 Dev at 2K resolution, with MIT license and ComfyUI support.
Moonshot AI closed a 2 billion dollar round at a 20 billion dollar valuation led by Meituan May 7. What the open-weights tier consolidation means for creators.
By mid-2026 ComfyUI hosts thousands of community workflows. Here are 12 that have crossed the production-quality threshold, with download links and the use case each one actually serves.
IBM released Granite 4.1 on April 29, 2026 with dense decoder-only 3B, 8B, and 30B Apache 2.0 LLMs that walk back the Granite 4.0 MoE bet. The 8B dense beats the 32B-A9B MoE on most benchmarks. 128K production context, 512K extension.
Zed Industries shipped Zed 1.0 on April 29 2026, a Rust-built, GPU-accelerated code editor with parallel agents, edit prediction, and the open Agent Client Protocol.
Mistral shipped Medium 3.5 as a 128B dense multimodal model under modified MIT, with cloud Vibe agents and Le Chat Work mode on day one.