NVIDIA Nemotron Diffusion: 3x Faster LLM Decoding

NVIDIA Nemotron Diffusion: 3x Faster LLM Decoding

NVIDIA released the Nemotron-Labs-Diffusion family on Hugging Face, an open-weights LLM that switches between autoregressive, diffusion, and self-speculation decoding for 2.7x to 3.3x throughput gains.

ByteDance Lance: 3B Open Model for Image and Video

ByteDance Lance: 3B Open Model for Image and Video

ByteDance Research released Lance, a 3B Apache 2.0 unified multimodal model that handles image and video generation, editing, and understanding in a single framework. Strong VBench and GenEval scores.

Zerostack: A Rust Coding Agent With 8MB RAM

Zerostack: A Rust Coding Agent With 8MB RAM

Zerostack is a pure Rust coding agent that launched May 16, 2026, running in 8MB of RAM compared to 300MB for JavaScript-based alternatives like Opencode.

IBM Granite Embedding R2: Open Multilingual RAG Models

IBM Granite Embedding R2: Open Multilingual RAG Models

IBM has released Granite Embedding Multilingual R2: a pair of Apache 2.0 embedding models with 32K context, 200+ languages, and a top MTEB score under 100M parameters. A drop-in swap for paid commercial embedding APIs.

VOID, BiRefNet, and Gemma 4 Are Now in ComfyUI

VOID, BiRefNet, and Gemma 4 Are Now in ComfyUI

Three significant open-source models arrived in ComfyUI on May 14, 2026: VOID for video object deletion, BiRefNet for image segmentation, and Gemma 4 for multimodal reasoning.

Free Weekly Newsletter

Stay ahead of Creative AI

Join creators getting the latest AI tools, model releases, and workflow tips delivered weekly.

No spam. Unsubscribe anytime.