Tencent Hy3: Open 295B Agentic Coding Model
Tencent open-sourced Hy3, a 295B mixture-of-experts model (21B active, 256K context) under Apache 2.0. It beats GLM-5.2 at roughly half the size and is free on OpenRouter until July 21.
Tencent open-sourced Hy3, a 295B mixture-of-experts model (21B active, 256K context) under Apache 2.0. It beats GLM-5.2 at roughly half the size and is free on OpenRouter until July 21.
Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter agentic coding model trained entirely on Chinese chips. The weights just landed on Hugging Face, and it beats GPT-5.5 on SWE-bench Pro.
Claude Fable 5 is generally available again as of July 1, 2026: Anthropic's most capable widely released model for reasoning and long-horizon agents. What it is, the pricing, and the refusal-and-fallback change every builder must handle.
Anthropic launched Claude Sonnet 5 on June 30, 2026 as the new default model for Free and Pro, positioned close to Opus 4.8 at a lower price. Here is what changed, the tokenizer catch that resets your cost math, and how to switch your workflow.
Anthropic launched Claude Tag, an always-on AI teammate that lives in Slack with persistent memory, ambient mode, and autonomous async tasks.
Workweave open-sourced Weave Router, a drop-in proxy that routes each prompt to the best model for Claude Code, Codex, and Cursor, claiming 40 to 70 percent lower costs.
Context engineering just got a dedicated workbench. On June 23, 2026, DAI Studio launched a free visual tool that lets you design, see, and version the exact context you feed into a large language model before it runs.
Sakana Fugu is a multi-agent system that behaves like one model: send a request to a single OpenAI-compatible API and it routes across frontier models for you. Here is what it costs, how it scores, and why orchestration matters for builders.
Anthropic released Claude Fable 5, its most capable model, on June 9, 2026, free on paid plans through June 22. How it compares to Opus 4.8.
Xiaomi MiMo-V2.5-Pro-UltraSpeed claims 1000 tokens per second decode on a 1T MoE. Open-weights FP4 checkpoint plus a 2-week free API trial.
New ICML 2026 research shows transformer models can share attention projections, achieving up to 96.9% KV cache reduction with minimal accuracy loss.
NVIDIA released Nemotron 3 Ultra on June 1 2026: a 550B mixture-of-experts model with 55B active parameters, open weights on Hugging Face, with 5x faster inference and 30% lower cost than Nemotron 2.
Developer Oscar Molnar installed a secondhand Tesla V100 SXM2 into his gaming PC alongside an RTX 4080, building a 32GB dual-GPU setup for under £200 total.
George Hotz argues AI agents frontload impressive progress but stall on polish, creating a golden era of AI-generated slop that creators need to plan for.
Qwen3.7-Max sets a new bar for non-hallucination rate among frontier agent models. Here is how it stacks up against Claude Opus 4.7, Gemini 3.1 Pro, and GPT-5.5 across the four reliability dimensions that decide which model goes into production.
Cohere released Command A+ under Apache 2.0 on May 21, 2026. The 218B sparse MoE runs on two H100s, with native citations and 48 languages.
NVIDIA released the Nemotron-Labs-Diffusion family on Hugging Face, an open-weights LLM that switches between autoregressive, diffusion, and self-speculation decoding for 2.7x to 3.3x throughput gains.
Google launched Gemini 3.5 Flash at I/O 2026, a Flash-tier model that beats 3.1 Pro on coding and agent benchmarks at 40 percent lower cost.
Alibaba pushed Qwen 3.7 Max Preview and Qwen 3.7 Plus Preview to Arena and Qwen Chat for testing. Max sits 13th overall on Arena Text and Plus is 16th on Vision.
DeepSeek-V4-Flash is the first local model competitive with frontier AI, making LLM activation steering practical for the first time. A guide for creators.
Civo Limited’s RelaxAI offers UK-sovereign LLM inference at £0.10 per million input tokens, with ISO 27001 certification and 100% UK data residency.
OpenAI enlisted outside legal counsel to explore action against Apple after the ChatGPT-Siri integration failed to generate expected subscription revenue.
OpenAI Codex is now in the ChatGPT iOS app, letting creators monitor projects, review diffs, and push code changes without being at a desk.
Richard Socher's stealth-mode AI lab Recursive exits with $650M at a $4.65B valuation, backed by GV, Greycroft, Nvidia, and AMD.