Deep Dive
Sep 12, 2026
Specific Labs published Real-SWE on September 12, 2026, a coding-agent benchmark built entirely from private production codebases. The top score is 38.8%, from Claude Fable 5.1 in Claude Code. The same generation posts 95% on SWE-bench Verified.
Deep Dive
Sep 10, 2026
Cognition released SWE-2 on 10 September 2026 and one number travelled: 50.0% on FrontierCode 1.1 Main, within a point of Fable 5.1, at 64% lower cost. Three rows below it, in the same table, is a number that did not.
News
May 26, 2026
Datacurves new DeepSWE benchmark ranks GPT-5.5 at 70% on 113 hand-written software-engineering tasks and exposes Claude Opus 4.6 and 4.7 retrieving git-history solutions on SWE-Bench Pro.
Video Generation
Apr 8, 2026
A model called HappyHorse-1.0 has taken the top spot on Artificial Analysis text-to-video leaderboard with an ELO rating of 1365, beating Seedance 2.0 and Kling 3.0 Pro.