Deep Dive
Sep 8, 2026
On September 8, 2026, European lab Desert Ant Labs published 18 small AI models that run entirely on a users device, with SDKs for Swift, Kotlin and JavaScript and a free tier covering the first 100,000 monthly active devices per platform.
Deep Dive
Sep 7, 2026
OpenBMB released MiniCPM5-2B on September 7, 2026, a 2.6B dense Apache-2.0 model that Artificial Analysis ranks #1 of 47 open-weights models at or under 4B parameters. The Q4_K_M quantization is 1.56 GB on disk.
Open Source
Sep 4, 2026
llama.cpp shipped version 0.4.0 on September 4, adding initial support for Qwen3.8-Flash-Next, NVIDIA Nemotron-3-Puzzle-75B-A9B, and video input for local multimodal models.
Deep Dive
Sep 4, 2026
Microsoft announced Project Zenith on September 4, 2026: a ready-to-code Windows 11 experience gated behind 64 GB of unified memory and 250 GB/s of bandwidth, so developers can run 30B+ models locally and unmetered.
Deep Dive
Sep 3, 2026
NVIDIA PAIR is a free, open-source Personal AI Router that turns the idle GPUs across your home or studio into one local inference endpoint. Here is how it works and when it beats the cloud.
Deep Dive
Sep 1, 2026
Perplexity launched Hybrid Compute for its Computer agent on Apple silicon Macs, splitting a single task between cloud and local models so confidential files never leave the device.
AI
Aug 25, 2026
Apple unveiled the M6 and M5 Ultra chips on August 25, 2026, and the headline is on-device AI: the M5 Ultra can run and fine-tune 100B-plus parameter models entirely on a Mac.
AI
Aug 11, 2026
Unsloth AI released Unsloth Desktop, a free open-source app that runs and fine-tunes LLMs and diffusion models locally, 2x faster with 70 percent less VRAM.
AI
Aug 9, 2026
Liquid AI released LFM2.5-2.6B, a 2.6-billion-parameter open-weights agentic model that runs entirely on-device, planning and calling tools with no cloud API cost.
Open Source
Jul 30, 2026
Thinking Machines released Inkling-Small on July 30, a 276B open-weight MoE with 12B active params that matches or beats the full Inkling on reasoning and coding at a third the token price.
Deep Dive
Jul 17, 2026
Thinking Machines released Inkling, a 975B-parameter open-weights Mixture-of-Experts model that reads text, images, and audio. You can self-host it or fine-tune it via Tinker.
Deep Dive
Jul 15, 2026
AMD's Lemonade 11.0 turns its local AI server into a full multimodal stack, adding text-to-speech and 3D generation alongside LLMs and image gen.
AI
Jul 14, 2026
Bonsai 27B is a new open-weights multimodal model that PrismML compressed to as little as 3.9 GB, small enough to run entirely on a phone or laptop with no cloud connection.
Open Source
Jul 13, 2026
Unsloth released NVFP4 quantized versions of Qwen3.6 that run up to 2.5x faster, with the 27B model fitting on a single 24GB GPU.
Open Source
Jul 10, 2026
A developer released Colibri, a pure-C engine that runs GLM-5.2 (744B MoE) on a 25GB-RAM machine with no GPU by streaming experts from an NVMe SSD.
Deep Dive
Jul 9, 2026
aria is a dependency-free native runtime that runs the full Stable Audio 3 text-to-music pipeline on ordinary GPUs, CPU-only laptops, and an 8GB Raspberry Pi 5, no Python required.
Deep Dive
Jul 6, 2026
AMD put a 128GB local AI workstation on retail shelves for $3,999. The Ryzen AI Halo runs models up to 200 billion parameters and undercuts Nvidia's DGX Spark by $700.
Deep Dive
Jun 26, 2026
Run a capable AI coding agent entirely on your own machine: one open-weight model, one GPU, no per-token bill, and no code leaving your network.
Deep Dive
Jun 10, 2026
Google DeepMind released DiffusionGemma on June 10, 2026, an Apache 2.0 open model that generates text up to 4x faster by denoising blocks of tokens in parallel, and it runs on a single RTX GPU.
Deep Dive
Jun 4, 2026
LM Studio released LM Link support for iPhone and iPad on June 4, 2026, through its official Locally mobile app.
AI Tools
Jun 1, 2026
NVIDIA shipped NemoClaw on June 1, 2026, a single-command installer for local AI agents on DGX Spark hardware, with multi-node clustering up to 512GB pooled memory.
AI Tools
Jun 1, 2026
WorkInProcess is a free browser-based image studio with AI upscaling and object removal. No uploads, no accounts, everything runs locally.
AI Tools
Jun 1, 2026
mistral.rs v0.8.2 delivers 3.5-5.5x faster MoE prefill on CUDA, fused decode kernels, and agentic tool-calling improvements for local LLM workflows.
AI
Jun 1, 2026
NVIDIA announced DGX Station for Windows at Computex 2026, putting a GB300 Grace Blackwell Ultra chip and 748GB of memory into a desktop that runs trillion-parameter models locally.