Desert Ant Labs: 18 On-Device AI Models for Creators

Desert Ant Labs: 18 On-Device AI Models for Creators

On September 8, 2026, European lab Desert Ant Labs published 18 small AI models that run entirely on a users device, with SDKs for Swift, Kotlin and JavaScript and a free tier covering the first 100,000 monthly active devices per platform.

MiniCPM5-2B: The 1.6GB Local Model That Runs Agents

MiniCPM5-2B: The 1.6GB Local Model That Runs Agents

OpenBMB released MiniCPM5-2B on September 7, 2026, a 2.6B dense Apache-2.0 model that Artificial Analysis ranks #1 of 47 open-weights models at or under 4B parameters. The Q4_K_M quantization is 1.56 GB on disk.

Project Zenith: 30B Local Coding Models on Windows

Project Zenith: 30B Local Coding Models on Windows

Microsoft announced Project Zenith on September 4, 2026: a ready-to-code Windows 11 experience gated behind 64 GB of unified memory and 250 GB/s of bandwidth, so developers can run 30B+ models locally and unmetered.

NVIDIA PAIR: Pool Every GPU for Local AI

NVIDIA PAIR: Pool Every GPU for Local AI

NVIDIA PAIR is a free, open-source Personal AI Router that turns the idle GPUs across your home or studio into one local inference endpoint. Here is how it works and when it beats the cloud.

Perplexity Hybrid Compute: Local-Cloud AI on Mac

Perplexity Hybrid Compute: Local-Cloud AI on Mac

Perplexity launched Hybrid Compute for its Computer agent on Apple silicon Macs, splitting a single task between cloud and local models so confidential files never leave the device.

Inkling-Small: Open 276B Model Beats Its Big Sibling

Inkling-Small: Open 276B Model Beats Its Big Sibling

Thinking Machines released Inkling-Small on July 30, a 276B open-weight MoE with 12B active params that matches or beats the full Inkling on reasoning and coding at a third the token price.

Inkling: Thinking Machines' Open Multimodal AI

Inkling: Thinking Machines' Open Multimodal AI

Thinking Machines released Inkling, a 975B-parameter open-weights Mixture-of-Experts model that reads text, images, and audio. You can self-host it or fine-tune it via Tinker.

Bonsai 27B Runs a 27B AI Model on Your Phone

Bonsai 27B Runs a 27B AI Model on Your Phone

Bonsai 27B is a new open-weights multimodal model that PrismML compressed to as little as 3.9 GB, small enough to run entirely on a phone or laptop with no cloud connection.

AMD Ryzen AI Halo: $3,999 Local AI Workstation

AMD Ryzen AI Halo: $3,999 Local AI Workstation

AMD put a 128GB local AI workstation on retail shelves for $3,999. The Ryzen AI Halo runs models up to 200 billion parameters and undercuts Nvidia's DGX Spark by $700.

DiffusionGemma: Google's 4x Faster Open Text Model

DiffusionGemma: Google's 4x Faster Open Text Model

Google DeepMind released DiffusionGemma on June 10, 2026, an Apache 2.0 open model that generates text up to 4x faster by denoising blocks of tokens in parallel, and it runs on a single RTX GPU.

Free Weekly Newsletter

Stay ahead of Creative AI

Join creators getting the latest AI tools, model releases, and workflow tips delivered weekly.

No spam. Unsubscribe anytime.