Gemma 4 12B: Encoder-Free Multimodal in 12 Billion Params
Google's new Gemma 4 12B drops separate vision and audio encoders, packing native video and speech understanding into a single 12B open-weights model that runs locally.
Google's new Gemma 4 12B drops separate vision and audio encoders, packing native video and speech understanding into a single 12B open-weights model that runs locally.
Google AI Edge Gallery arrives on macOS with Gemma 4 12B support, bringing local AI model testing to Mac users for the first time.
Ideogram released its first open-weight text-to-image model on June 3, 2026: a 9.3B parameter Diffusion Transformer with JSON-structured prompting, in-image text rendering, and day-zero ComfyUI support.
TinyFish Bigset launched June 2 as an AGPL-3.0 open-source multi-agent system that turns a plain-English sentence into a structured dataset pulled from the live web, then refreshes it on a schedule.
NVIDIA forms Cosmos Coalition with Runway, Black Forest Labs, and LTX to build open video generation infrastructure.
JetBrains releases Mellum2 Thinking under Apache 2.0, an open-weights coding model with chain-of-thought reasoning.
Fizgig v1.2.4 makes full Flux 2 Klein 9B LoRA training possible on 16GB GPUs using fp8 Base DiT at 9.6GB VRAM. The free, open-source studio includes training presets, repair tools, and profiler.
NVIDIA pushed new PiD checkpoints June 2 with a FLUX.2 color-fix variant plus Qwen-Image support, all on Apache 2.0 for direct 4K decode in ComfyUI.
Deploy H Company's Apache 2.0 Holo 3.1 4B locally with vLLM on a 12GB consumer GPU, hook it into a desktop-agent workflow, and benchmark step latency and cost against Claude Computer Use and OpenAI Operator.
H Company released Holo 3.1, an Apache 2.0 computer-use agent family with quantized weights, mobile control, and 79.3% AndroidWorld score.
mistral.rs v0.8.2 delivers 3.5-5.5x faster MoE prefill on CUDA, fused decode kernels, and agentic tool-calling improvements for local LLM workflows.
NVIDIA released Nemotron 3 Ultra on June 1 2026: a 550B mixture-of-experts model with 55B active parameters, open weights on Hugging Face, with 5x faster inference and 30% lower cost than Nemotron 2.
StepFun ships Step 3.7 Flash: open-weights 201B MoE VLM with 256K context, native video input, and FP8/NVFP4 builds for local deployment.
OpenMOSS published the MOSS-Audio technical report on June 1, 2026, documenting four open-source audio-language models that achieve benchmark scores rivaling systems three to four times their size.
ByteDance releases Bernini-R, an open-source video editing model that handles object insertion, removal, and replacement in existing footage using text prompts.
PewDiePie open-sourced Odysseus, a self-hosted AI workspace with chat, agents, deep research, and 270+ model serving. MIT license, no telemetry.
NVIDIA Cosmos 3 ranks first among open-source models on the Artificial Analysis Text-to-Image leaderboard. The 64B-parameter model ships with image-to-video capability and open commercial-use weights.
Pallaidium turns Blender's Video Sequence Editor into a complete AI movie studio. The May 31 release adds Blender 5.2 support and a redesigned plugin architecture.
A developer used Stability AI Stable Audio 3 Medium to generate 15,834 free audio samples: 10,359 drum one-shots and 5,475 pitched instrument recordings, available for immediate download.
Llama Studio v0.2.0 is a lightweight web interface for managing multiple llama-server sessions, with multi-GPU tensor splitting, shell-script configs, and auto-load snapshots.
Stable Audio Studio is a new open-source desktop application that runs Stability AI's Stable Audio Open 1.0 model directly on your machine. Released to GitHub in May 2026, the project gives creators a full audio generation environment with no account or subscription.
A new open-source ComfyUI custom node called TextMakerPro brings a layer-based text and layout editor into Stable Diffusion workflows, letting creators design stylized text compositions without leaving ComfyUI.
Liquid AI dropped LFM2.5-8B-A1B on Hugging Face on May 28, the first reasoning-tuned MoE in the LFM2.5 family with 8.3B params, 1.5B active per token, and built-in tool calling.
Baidu open-sourced NAVA, a 6.3B parameter joint audio-video model that generates 720p video with synced dual-channel audio in a single pass.