GPT-5.6 Sol, Terra, Luna Land on Amazon Bedrock
OpenAI's GPT-5.6 family is now on Amazon Bedrock, giving builders three model tiers, Sol, Terra, and Luna, inside the AWS stack they already use.
OpenAI's GPT-5.6 family is now on Amazon Bedrock, giving builders three model tiers, Sol, Terra, and Luna, inside the AWS stack they already use.
Claude Opus 5 launched July 24, 2026 at the same $5/$25 pricing as Opus 4.8, but Anthropic says it more than doubles 4.8's performance and undercuts Fable 5 on cost per completed task.
Echo, a new public-alpha endpoint from Tracer, pools open-weight models like GLM-5.2 and Kimi behind one OpenAI-compatible API, promising Claude-class output at roughly a third of the cost.
NvChat is a free, open-source Windows app that gives you a native desktop chat interface for NVIDIA free hosted LLM API, with access to more than 100 models including vision and reasoning models.
Anthropic updated Claude voice mode to run on Opus and Sonnet, not just Haiku, and to act inside Gmail, Calendar, Slack, Canva, and Notion by voice. Here is what changed and how to use it.
Upstage released Solar Open 2, a 250-billion-parameter open-weight MoE built for long-horizon agentic coding, tool calling, and document work.
Alibaba's Qwen team announced Qwen3.8, its next flagship model, saying it will go open-weight soon. As of July 19, 2026 there are no weights or benchmarks yet.
Thinking Machines released Inkling, a 975B-parameter open-weights Mixture-of-Experts model that reads text, images, and audio. You can self-host it or fine-tune it via Tinker.
Moonshot AI has officially launched Kimi K3, the world's first open 3T-class model: a 2.8T-parameter MoE with a 1M-token context window and open weights arriving July 27.
Google's Gemini 3.5 Pro is reportedly targeting a July 17 general-availability launch, according to leaked launch plans and third-party reporting rather than any official Google post.
Unsloth released NVFP4 quantized versions of Qwen3.6 that run up to 2.5x faster, with the 27B model fitting on a single 24GB GPU.
A developer released Colibri, a pure-C engine that runs GLM-5.2 (744B MoE) on a 25GB-RAM machine with no GPU by streaming experts from an NVMe SSD.
Subagentmaxxing is a new open-source CLI that runs Codex, Cursor, and Grok as subagents under Claude Code, keeping flagship planning while cutting execution costs up to 25 times.
Ollama raised $88M and now serves 8.9M developers. Here is how to run open AI models locally, when to reach for Ollama Cloud, and a full getting-started workflow.
Mistral added version control for prompts and skills to Mistral Studio on July 9, 2026, with immutable versions, rollback, ownership tracking, and audit logs.
OpenAI launched GPT-Live on July 8, 2026, a pair of full-duplex voice models that listen and speak at the same time, replacing Advanced Voice Mode for every ChatGPT user.
Hugging Face and vLLM shipped a native-speed transformers backend on July 8, 2026, serving almost any Hugging Face model at full vLLM speed with a single flag.
Tencent open-sourced Hy3, a 295B mixture-of-experts model (21B active, 256K context) under Apache 2.0. It beats GLM-5.2 at roughly half the size and is free on OpenRouter until July 21.
Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter agentic coding model trained entirely on Chinese chips. The weights just landed on Hugging Face, and it beats GPT-5.5 on SWE-bench Pro.
Claude Fable 5 is generally available again as of July 1, 2026: Anthropic's most capable widely released model for reasoning and long-horizon agents. What it is, the pricing, and the refusal-and-fallback change every builder must handle.
Anthropic launched Claude Sonnet 5 on June 30, 2026 as the new default model for Free and Pro, positioned close to Opus 4.8 at a lower price. Here is what changed, the tokenizer catch that resets your cost math, and how to switch your workflow.
Anthropic launched Claude Tag, an always-on AI teammate that lives in Slack with persistent memory, ambient mode, and autonomous async tasks.
Workweave open-sourced Weave Router, a drop-in proxy that routes each prompt to the best model for Claude Code, Codex, and Cursor, claiming 40 to 70 percent lower costs.
Context engineering just got a dedicated workbench. On June 23, 2026, DAI Studio launched a free visual tool that lets you design, see, and version the exact context you feed into a large language model before it runs.