Google Cloud API Gateway Now Routes to Any LLM
Google shipped model routing for Cloud API Gateway on August 4, 2026: one OpenAI-compatible endpoint that forwards each request to Gemini, Claude, or an OpenAI GPT model.
Google shipped model routing for Cloud API Gateway on August 4, 2026: one OpenAI-compatible endpoint that forwards each request to Gemini, Claude, or an OpenAI GPT model.
Cloudflare shipped three inference optimizations that make Kimi K2.6 and GLM 5.2 roughly 30% cheaper to run on Workers AI, and its benchmark tables show the answers do not change.
Alibaba announced Qwen3.8-Max on August 3, 2026: a 2.4 trillion-parameter coding flagship with a 1 million-token context and a promised open-weights release.
DeepSeek shipped the official production build of V4 Flash on July 31, 2026: an open-weight, MIT-licensed 284B-parameter MoE model with a 1M token context, native OpenAI Responses API and Codex support, and API output at $0.28 per million tokens.
Thinking Machines released Inkling-Small on July 30, a 276B open-weight MoE with 12B active params that matches or beats the full Inkling on reasoning and coding at a third the token price.
OpenAI cut API prices on its two cheaper GPT-5.6 tiers on July 30, 2026, dropping Luna 80% and Terra 20% while leaving the top Sol tier unchanged.
TurboFieldfare is an open-source Swift and Metal engine that runs Google Gemma 4 26B in about 2GB of RAM on any M-series Mac by streaming experts from SSD.
OpenAI's GPT-5.6 family is now on Amazon Bedrock, giving builders three model tiers, Sol, Terra, and Luna, inside the AWS stack they already use.
Claude Opus 5 launched July 24, 2026 at the same $5/$25 pricing as Opus 4.8, but Anthropic says it more than doubles 4.8's performance and undercuts Fable 5 on cost per completed task.
Echo, a new public-alpha endpoint from Tracer, pools open-weight models like GLM-5.2 and Kimi behind one OpenAI-compatible API, promising Claude-class output at roughly a third of the cost.
NvChat is a free, open-source Windows app that gives you a native desktop chat interface for NVIDIA free hosted LLM API, with access to more than 100 models including vision and reasoning models.
Anthropic updated Claude voice mode to run on Opus and Sonnet, not just Haiku, and to act inside Gmail, Calendar, Slack, Canva, and Notion by voice. Here is what changed and how to use it.
Upstage released Solar Open 2, a 250-billion-parameter open-weight MoE built for long-horizon agentic coding, tool calling, and document work.
Alibaba's Qwen team announced Qwen3.8, its next flagship model, saying it will go open-weight soon. As of July 19, 2026 there are no weights or benchmarks yet.
Thinking Machines released Inkling, a 975B-parameter open-weights Mixture-of-Experts model that reads text, images, and audio. You can self-host it or fine-tune it via Tinker.
Moonshot AI has officially launched Kimi K3, the world's first open 3T-class model: a 2.8T-parameter MoE with a 1M-token context window and open weights arriving July 27.
Google's Gemini 3.5 Pro is reportedly targeting a July 17 general-availability launch, according to leaked launch plans and third-party reporting rather than any official Google post.
Unsloth released NVFP4 quantized versions of Qwen3.6 that run up to 2.5x faster, with the 27B model fitting on a single 24GB GPU.
A developer released Colibri, a pure-C engine that runs GLM-5.2 (744B MoE) on a 25GB-RAM machine with no GPU by streaming experts from an NVMe SSD.
Subagentmaxxing is a new open-source CLI that runs Codex, Cursor, and Grok as subagents under Claude Code, keeping flagship planning while cutting execution costs up to 25 times.
Ollama raised $88M and now serves 8.9M developers. Here is how to run open AI models locally, when to reach for Ollama Cloud, and a full getting-started workflow.
Mistral added version control for prompts and skills to Mistral Studio on July 9, 2026, with immutable versions, rollback, ownership tracking, and audit logs.
OpenAI launched GPT-Live on July 8, 2026, a pair of full-duplex voice models that listen and speak at the same time, replacing Advanced Voice Mode for every ChatGPT user.
Hugging Face and vLLM shipped a native-speed transformers backend on July 8, 2026, serving almost any Hugging Face model at full vLLM speed with a single flag.