DeepSeek V4.1-Flash Retires V4-Pro and Cuts Prices
DeepSeek shipped V4.1-Flash on 10 September: 552B multimodal MoE, MIT open weights, cache-hit prices down 57%, and every V4-Pro API call rerouted to it from 14 September.
DeepSeek shipped V4.1-Flash on 10 September: 552B multimodal MoE, MIT open weights, cache-hit prices down 57%, and every V4-Pro API call rerouted to it from 14 September.
DeepSeek added image understanding to its cheapest fast model on August 21, 2026, billing images at standard V4-Flash rates with a free Files API.
Google shipped Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash, with double-digit coding and agent benchmark gains and half-price intro pricing.
OpenAI's Ultrafast tier runs GPT-5.6 Sol at up to 750 output tokens per second, about 14x faster than Standard, powered by Cerebras wafer-scale hardware.
Google shipped Gemini 3.6 Flash on July 21, 2026: $7.50 output pricing (down 17%), 17% fewer output tokens, and stronger agentic coding. What it means for builders.
Mozilla released its inaugural State of Open Source AI report on July 14, 2026. Open-weight models have closed the capability gap with closed frontier systems to just 3.3 percent, while inference costs fell roughly 50 times in three years.
Cognition SWE-1.7 lands within a few points of GPT-5.5 and Claude Opus 4.8 on coding benchmarks, at a fraction of the cost. It is live in Devin at 1000 tokens per second.
Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter agentic coding model trained entirely on Chinese chips. The weights just landed on Hugging Face, and it beats GPT-5.5 on SWE-bench Pro.
Hugging Face benchmarked LoRA against newer PEFT methods like OFT, which beat it on image generation while using less memory. The default is not always the best fit.
Google is building a Skills Marketplace for Gemini Enterprise, a hub where teams pick predefined AI skills instead of configuring each capability by hand.
Anthropic disabled Claude Fable 5 and Mythos 5 for every customer worldwide after a US government export-control directive on June 12, 2026.
Moonshot released Kimi K2.7-Code on June 12, an open-weights coding model that beats Claude Opus 4.8 on MCPMark tool use at far lower cost.
Anthropic apologized on June 11 for a hidden Claude Fable 5 guardrail that quietly degraded results instead of refusing them, and said it will make the safeguard visible.
Apple WWDC 2026 unveiled five new Foundation Models, a rebuilt Siri AI, Image Playground, and Spatial Reframing, built with Google Gemini collaboration.
Xiaomi MiMo-V2.5-Pro-UltraSpeed claims 1000 tokens per second decode on a 1T MoE. Open-weights FP4 checkpoint plus a 2-week free API trial.
Microsoft used its Build 2026 keynote on June 2 to ship MAI-Code-1-Flash, an in-house coding model that is live today inside the GitHub Copilot model picker for Free, Pro, Pro+, and Max tiers.
Microsoft's MAI-Thinking-1, a 35B sparse MoE with 256k context, tops Claude Sonnet 4.6 in blind evaluations and matches Claude Opus 4.6 on SWE-Bench Pro.
Nvidia is bringing Cosmos, Nemotron, GR00T, and Ising under OpenMDW-1.1. Here is what the unified AI model license means for creative AI developers.
Xiaomi cut MiMo-v2.5 API prices up to 99% effective May 27 across chat, image, audio, video understanding, and TTS, with OpenAI and Anthropic compatible endpoints.
PrismML released Bonsai Image 4B with 1-bit and ternary checkpoints under Apache 2.0. The model retains 95% of FLUX.2 Klein 4B quality at 6.4x smaller size and runs directly on iPhone.
Alibaba released Qwen3.6-27B on April 22, 2026. The 27B dense open-weight model beats its larger 35B-A3B sibling on coding, agent, and vision benchmarks.
Alibaba released Qwen3.6-Max-Preview on April 20, 2026, topping six programming benchmarks including SWE-benchPro. Here is what it means for creators building AI workflows.
Meta launched Muse Spark on April 8, the first AI model built by its new Superintelligence Labs division. Unlike every Llama release before it, Muse Spark is proprietary.
A model called HappyHorse-1.0 has taken the top spot on Artificial Analysis text-to-video leaderboard with an ELO rating of 1365, beating Seedance 2.0 and Kling 3.0 Pro.