Six Open Jev Clones, Four Different Jev Scores

Six Open Jev Clones, Four Different Jev Scores

TypeSafe opened access to Jev on 15 September. Within 48 hours six open reproductions had shipped, each claiming to match it. Four independent evaluations put Jev's own Banking77 accuracy at 87.0%, 83.2%, 77.8% and 76.3%.

Agnes 3.0 Flash: The Open Weights Aren't the Leaderboard

Agnes 3.0 Flash: The Open Weights Aren't the Leaderboard

Agnes AI published Agnes-3.0-Flash Preview, a 33B hybrid-attention multimodal model with a 262,144-token context window, under Apache 2.0 on 11 September 2026. The downloadable weights are a different checkpoint from the Agnes 3.0 Flash listed on Artificial Analysis.

Cohere North Small Translate Beats DeepL, With a Catch

Cohere North Small Translate Beats DeepL, With a Catch

Cohere released North Small Translate 1.0 on 10 September 2026, a 218B model scoring 83.60 on WMT26 against DeepL NextGen at 81.37. The weights are CC BY-NC 4.0, so you cannot use them in anything you sell.

MiniCPM5-2B: The 1.6GB Local Model That Runs Agents

MiniCPM5-2B: The 1.6GB Local Model That Runs Agents

OpenBMB released MiniCPM5-2B on September 7, 2026, a 2.6B dense Apache-2.0 model that Artificial Analysis ranks #1 of 47 open-weights models at or under 4B parameters. The Q4_K_M quantization is 1.56 GB on disk.

OpenAI Begins Rolling Out GPT-6 Astra

OpenAI Begins Rolling Out GPT-6 Astra

OpenAI began rolling out GPT-6 Astra on September 3, its most capable model for computer use, browsing, and software engineering, with paid ChatGPT tiers and the API to follow in the coming days.

Funes: The Coding-Agent Memory You Own

Funes: The Coding-Agent Memory You Own

Hugging Face's open-source Funes gives coding agents like Claude Code and Codex a persistent memory that lives on your machine and travels between tools. Here is how the agent-memory-you-own pattern compares to vendor-locked and knowledge-graph memory.

K2 Horizon: World's Largest Fully Open AI Models

K2 Horizon: World's Largest Fully Open AI Models

IFM released K2 Horizon on September 3: six fully open models from 0.9B to 375B, shipped with weights, code, and training data under Apache 2.0. Here is how to download and run them.

Anthropic Open-Sources Claude Commerce Agents

Anthropic Open-Sources Claude Commerce Agents

Anthropic open-sourced Claude Commerce Agents on September 2, 2026, a forkable blueprint for shopping and merchant agents. Here is the fork-to-shipped builder path, and how the open blueprint compares to a raw SDK build and closed vertical platforms.

Meta Muse Spark 1.3 and the Efficiency Turn in Coding AI

Meta Muse Spark 1.3 and the Efficiency Turn in Coding AI

Meta shipped Muse Spark 1.3 with a pitch built on efficiency: roughly 20 percent fewer tool calls and 25 percent fewer tokens per coding task. Here is why tokens and tool calls per completed task, not raw benchmark score, are the metric that actually decides what an agent costs to run.

Celeris-1 Magnus: A Diffusion Model Built for Agents

Celeris-1 Magnus: A Diffusion Model Built for Agents

Celeris-1 Magnus is a hybrid diffusion language model built for agentic work that the startup says runs faster than autoregressive models on tool-heavy tasks. It exposes an OpenAI-compatible API available today.

Claude Fable 5.1 Cuts Cache Reads by 75 Percent

Claude Fable 5.1 Cuts Cache Reads by 75 Percent

Anthropic released Claude Fable 5.1 on September 1, 2026. Headline pricing did not move, but cache reads dropped from $1.00 to $0.25 per million tokens, a 75 percent cut worth up to 45 percent on agentic workloads.

Mercury 2.5 Preview: Fast, Cheap Diffusion LLM

Mercury 2.5 Preview: Fast, Cheap Diffusion LLM

Inception released Mercury 2.5 Preview on August 31, 2026, a diffusion LLM that generates tokens in parallel, hitting up to 1,107 tokens per second at cost-optimized pricing on OpenRouter.

Tencent Hy4 Preview: 770B Open-Weights Model

Tencent Hy4 Preview: 770B Open-Weights Model

Tencent has open-sourced Hy4 preview, a 770B mixture-of-experts model with a 1M-token context window under Apache 2.0, aimed at coding, data analysis, and builder workflows.

Frontier AI API Prices Dropped: What It Costs Now

Frontier AI API Prices Dropped: What It Costs Now

OpenAI cut developer pricing for GPT-5.6 Sol on August 21, 2026. Input fell from $5 to $4 per million tokens and output from $30 to $20, putting it below Claude Opus 5 on both sides.

DeepSeek V4 Pro Ships: 1.6T Flagship, Open Weights

DeepSeek V4 Pro Ships: 1.6T Flagship, Open Weights

DeepSeek pushed its flagship V4 Pro to general availability on August 12, 2026, ending a preview that ran nearly four months. The GA build is a 1.6 trillion parameter MoE with a 1M-token context window.

Grok 4.6: Frontier Agentic Coding at a Lower Cost

Grok 4.6: Frontier Agentic Coding at a Lower Cost

Grok 4.6 hits 61 on the Artificial Analysis Intelligence Index at $2/$6 pricing and finishes agent tasks in roughly half the turns of premium rivals. Full benchmarks and a creator workflow.

Free Weekly Newsletter

Stay ahead of Creative AI

Join creators getting the latest AI tools, model releases, and workflow tips delivered weekly.

No spam. Unsubscribe anytime.