Inkling: Thinking Machines' Open Multimodal AI

Inkling: Thinking Machines' Open Multimodal AI

Thinking Machines released Inkling, a 975B-parameter open-weights Mixture-of-Experts model that reads text, images, and audio. You can self-host it or fine-tune it via Tinker.

Bonsai 27B Runs a 27B AI Model on Your Phone

Bonsai 27B Runs a 27B AI Model on Your Phone

Bonsai 27B is a new open-weights multimodal model that PrismML compressed to as little as 3.9 GB, small enough to run entirely on a phone or laptop with no cloud connection.

AMD Ryzen AI Halo: $3,999 Local AI Workstation

AMD Ryzen AI Halo: $3,999 Local AI Workstation

AMD put a 128GB local AI workstation on retail shelves for $3,999. The Ryzen AI Halo runs models up to 200 billion parameters and undercuts Nvidia's DGX Spark by $700.

DiffusionGemma: Google's 4x Faster Open Text Model

DiffusionGemma: Google's 4x Faster Open Text Model

Google DeepMind released DiffusionGemma on June 10, 2026, an Apache 2.0 open model that generates text up to 4x faster by denoising blocks of tokens in parallel, and it runs on a single RTX GPU.

Gemma 4 2B: Local Tool Calling and Code Review

Gemma 4 2B: Local Tool Calling and Code Review

Google's Gemma 4 E2B proves that a 2B parameter model can handle structured JSON output, tool calling, reasoning traces, and real code review -- all running locally at zero API cost.

llama.cpp Adds Audio Support With Qwen3-Omni

llama.cpp Adds Audio Support With Qwen3-Omni

llama.cpp release b8769 adds audio multimodal support for Qwen3-Omni and Qwen3-ASR models, bringing local speech recognition and audio understanding to consumer hardware.

Free Weekly Newsletter

Stay ahead of Creative AI

Join creators getting the latest AI tools, model releases, and workflow tips delivered weekly.

No spam. Unsubscribe anytime.