Deep Dive
Sep 19, 2026
TypeSafe opened access to Jev on 15 September. Within 48 hours six open reproductions had shipped, each claiming to match it. Four independent evaluations put Jev's own Banking77 accuracy at 87.0%, 83.2%, 77.8% and 76.3%.
Deep Dive
Sep 11, 2026
Agnes AI published Agnes-3.0-Flash Preview, a 33B hybrid-attention multimodal model with a 262,144-token context window, under Apache 2.0 on 11 September 2026. The downloadable weights are a different checkpoint from the Agnes 3.0 Flash listed on Artificial Analysis.
Deep Dive
Sep 10, 2026
Cohere released North Small Translate 1.0 on 10 September 2026, a 218B model scoring 83.60 on WMT26 against DeepL NextGen at 81.37. The weights are CC BY-NC 4.0, so you cannot use them in anything you sell.
Deep Dive
Sep 9, 2026
DeepSeek shipped V4.1-Flash on 10 September: 552B multimodal MoE, MIT open weights, cache-hit prices down 57%, and every V4-Pro API call rerouted to it from 14 September.
Deep Dive
Sep 7, 2026
OpenBMB released MiniCPM5-2B on September 7, 2026, a 2.6B dense Apache-2.0 model that Artificial Analysis ranks #1 of 47 open-weights models at or under 4B parameters. The Q4_K_M quantization is 1.56 GB on disk.
AI
Sep 4, 2026
H Company released NeoMME, an Apache 2.0 family of multimodal and multilingual encoders for visual document retrieval and RAG, in 260M and 800M sizes.
LLM
Sep 3, 2026
OpenAI began rolling out GPT-6 Astra on September 3, its most capable model for computer use, browsing, and software engineering, with paid ChatGPT tiers and the API to follow in the coming days.
Deep Dive
Sep 3, 2026
Hugging Face's open-source Funes gives coding agents like Claude Code and Codex a persistent memory that lives on your machine and travels between tools. Here is how the agent-memory-you-own pattern compares to vendor-locked and knowledge-graph memory.
Open Source
Sep 3, 2026
IFM released K2 Horizon on September 3: six fully open models from 0.9B to 375B, shipped with weights, code, and training data under Apache 2.0. Here is how to download and run them.
Deep Dive
Sep 2, 2026
Anthropic open-sourced Claude Commerce Agents on September 2, 2026, a forkable blueprint for shopping and merchant agents. Here is the fork-to-shipped builder path, and how the open blueprint compares to a raw SDK build and closed vertical platforms.
Deep Dive
Sep 2, 2026
Meta shipped Muse Spark 1.3 with a pitch built on efficiency: roughly 20 percent fewer tool calls and 25 percent fewer tokens per coding task. Here is why tokens and tool calls per completed task, not raw benchmark score, are the metric that actually decides what an agent costs to run.
Deep Dive
Sep 1, 2026
Mercury 2.5 and Celeris-1 Magnus are diffusion LLMs that decode tokens in parallel to cut agent latency. Here is how they compare and when to swap.
LLM
Sep 1, 2026
Celeris-1 Magnus is a hybrid diffusion language model built for agentic work that the startup says runs faster than autoregressive models on tool-heavy tasks. It exposes an OpenAI-compatible API available today.
Deep Dive
Sep 1, 2026
Anthropic released Claude Fable 5.1 on September 1, 2026. Headline pricing did not move, but cache reads dropped from $1.00 to $0.25 per million tokens, a 75 percent cut worth up to 45 percent on agentic workloads.
LLM
Aug 31, 2026
Inception released Mercury 2.5 Preview on August 31, 2026, a diffusion LLM that generates tokens in parallel, hitting up to 1,107 tokens per second at cost-optimized pricing on OpenRouter.
Deep Dive
Aug 28, 2026
Tencent has open-sourced Hy4 preview, a 770B mixture-of-experts model with a 1M-token context window under Apache 2.0, aimed at coding, data analysis, and builder workflows.
AI
Aug 26, 2026
Alibaba released Qwen3.8-Flash-Next as open weights on Hugging Face and ModelScope, the first public preview of its next-generation Qwen4 architecture.
AI
Aug 26, 2026
Z.ai officially launched GLM-5.3-Flash, the stealth Ox Alpha model that topped OpenRouter, as an MIT-licensed open-weight coding model.
Deep Dive
Aug 21, 2026
OpenAI cut developer pricing for GPT-5.6 Sol on August 21, 2026. Input fell from $5 to $4 per million tokens and output from $30 to $20, putting it below Claude Opus 5 on both sides.
AI
Aug 12, 2026
DeepSeek pushed its flagship V4 Pro to general availability on August 12, 2026, ending a preview that ran nearly four months. The GA build is a 1.6 trillion parameter MoE with a 1M-token context window.
Deep Dive
Aug 12, 2026
Grok 4.6 hits 61 on the Artificial Analysis Intelligence Index at $2/$6 pricing and finishes agent tasks in roughly half the turns of premium rivals. Full benchmarks and a creator workflow.
Deep Dive
Aug 11, 2026
OpenAI launched a ChatGPT desktop app for Linux on August 11, 2026, bringing ChatGPT, Work, and the Codex coding agent to Ubuntu, Debian, and Fedora.
Deep Dive
Aug 8, 2026
Which local fine-tuning tool should you install? Soup, Unsloth, Axolotl, and LLaMA-Factory compared on the constraints that actually decide it.
Deep Dive
Aug 6, 2026
OpenAI gave free ChatGPT users unlimited text chats and moved them to GPT-5.6 Luna, while Plus and Pro get the Sol flagship and a new thinking slider.