NaiveAI, a Beijing startup, released Naive-N0.5-Flash on 27 September 2026: an open-weight coding model with 309B total parameters, 15.5B active, a native 1M-token context window and an MIT license. The launch chart puts it 6.4 points behind DeepSeek V4.1 Flash on DeepSWE v1.1 (67.8 against 74.2). The two scores come from different agent harnesses, though. Measured the same way, inside Claude Code, the gap is 2 points.

That changes the buying question. NaiveAI has announced API pricing of $0.10 per million input tokens and $0.40 per million output, below DeepSeek's off-peak output rate. Which model is cheaper for your coding agent comes down to one number most people never check: your cache-hit rate. Below we compare the numbers side by side on matched harnesses, cost a real agent task on all three models, and do the hardware maths for self-hosting. NaiveAI's technical blog supplies the rest of the detail.

What NaiveAI Released

The release is three repositories on Hugging Face and a code repo on GitHub, all created on 27 September. The full-precision checkpoint is 617.8 GB. The FP8 checkpoint is 315.1 GB and is the one the model card's quick start loads. A 1.3 GB draft model (652.8M parameters) handles speculative decoding.

The model is not trained from scratch. It builds on Xiaomi's open MiMo-V2.5 base model, and NaiveAI replaced its global-attention layers with DeepSeek Sparse Attention (DSA). The result has 48 layers: 39 use sliding-window attention over 128 tokens, and 9 use DSA, which picks the top 2,048 tokens from the full history. No layer uses full attention. Continued training ran 3.25T tokens (50B of indexer warmup, 3T of sparse-attention training and 200B of learning-rate decay), all at the native 1M context.

The model card makes three claims that matter to a builder:

  • Coding and AI R&D focus. Every benchmark in the launch charts is either repository-level software engineering or machine-learning research work.
  • Fast inference. NaiveRT, the company's serving stack, reaches 50 tokens per second per user in Standard mode and up to 2,000 in Ultrafast mode.
  • Cheap API. $0.10 input, $0.40 output and $0.01 for cache reads, per million tokens. The card says API access "will also be provided". The endpoint at api.naive.ai was answering with an authentication error when we checked, so the service exists, but we could not confirm open self-serve sign-up.
3D render of a 309B parameter block with a 15.5B active subset highlighted
309B total parameters, 15.5B active per token.

The Benchmark Chart Mixes Harnesses

NaiveAI's evaluation notes say that, unless noted otherwise, Naive-N0.5-Flash was tested in Claude Code 2.1.207 with only basic file and Bash tools. The competitor scores are copied from each vendor's own report. For DeepSeek V4.1 Flash, that report is the DeepSeek model card, which ran DeepSWE in the mini-SWE harness and Terminal-Bench in DeepSeek's own harness. Those are that model's two best scaffolds.

The same DeepSeek card publishes a per-scaffold table, and it includes Claude Code. That gives us a matched comparison:

BenchmarkNaive-N0.5-Flash (Claude Code)DeepSeek V4.1 Flash (Claude Code)DeepSeek V4.1 Flash (as charted)
DeepSWE v1.167.869.874.2 (mini-SWE)
Terminal-Bench 2.186.788.090.6 (DeepSeek Harness)

On matched harnesses the lead shrinks from 6.4 points to 2.0 on DeepSWE, and from 3.9 to 1.3 on Terminal-Bench 2.1. DeepSeek still wins both, but by a margin you would struggle to feel on a single repository. One caveat: DeepSeek's per-scaffold numbers were run at maximum reasoning effort, and its card does not state which Claude Code version it used.

Harness choice moves one model more than the gap between these two models. DeepSeek V4.1 Flash alone spans 65.5 (OpenCode) to 74.2 (mini-SWE) on DeepSWE, an 8.7-point spread. We made the same point when DeepSeek shipped, in our V4.1 Flash deep dive. A launch chart that mixes scaffolds is ranking harnesses as much as models.

3D bars showing DeepSWE scores 67.8, 69.8 and 74.2
DeepSWE v1.1: Naive 67.8 and DeepSeek 69.8 in Claude Code, DeepSeek 74.2 in mini-SWE.

Head-to-Head: Naive vs DeepSeek vs GLM-5.3-Flash

The obvious third comparison is Z.ai's GLM-5.3-Flash, another MIT-licensed coding model with a 1M context. Its card says Terminal-Bench 2.1 ran in the same Claude Code 2.1.207 that NaiveAI used, which makes that row the cleanest three-way comparison available. Its DeepSWE score came from mini-SWE, so that row is not matched.

Naive-N0.5-FlashDeepSeek V4.1 FlashGLM-5.3-Flash
Total / active parameters309B / 15.5B552B / 8B-16B320B / 18B
LicenseMITMITMIT
Context1M1M1M
Weights on Hugging Face315.1 GB (FP8)510.3 GB328.4 GB
Terminal-Bench 2.1, Claude Code86.788.084.3
DeepSWE v1.167.8 (Claude Code)69.8 (Claude Code)63.4 (mini-SWE)
NL2Repo-Bench71.964.0not reported
ProgramBench (Almost@1)17.520.3not reported
API input / output per 1M$0.10 / $0.40 (announced)$0.15 / $0.60 off-peak, $0.30 / $1.20 peak$0.15 / $0.50
Cache read per 1M$0.01$0.003 off-peak, $0.006 peak$0.03

Naive-N0.5-Flash has one clear win: NL2Repo-Bench, which asks a model to build a whole repository from a natural-language spec. It scores 71.9 against DeepSeek's 64.0, and DeepSeek's number was measured in its own harness, its best case. On ProgramBench, which rebuilds programs from executables and documentation, DeepSeek leads 20.3 to 17.5. Naive also reports 32.4 on Agent's Last Exam against DeepSeek's self-reported 31.8, though the two used different scaffolds. Our earlier three-way open coding comparison covers GLM-5.3-Flash's own trade-offs.

What a Coding Agent Task Costs on Each

Agent sessions re-send the same repository context turn after turn, so most input tokens are cache reads. That is where the price tables diverge. DeepSeek charges $0.003 per million cached tokens off-peak on its pricing page, under a third of NaiveAI's $0.01. Z.ai charges $0.03 for GLM-5.3-Flash. NaiveAI is cheapest on uncached input and on output.

Take a typical agent task: 3M input tokens at a 90% cache-hit rate, plus 60K output tokens.

ModelCost per taskPer 100 tasks
Naive-N0.5-Flash (announced rates)$0.081$8.10
DeepSeek V4.1 Flash, off-peak$0.089$8.91
GLM-5.3-Flash$0.156$15.60
DeepSeek V4.1 Flash, peak$0.178$17.82

Now stretch it to a long session: 10M input at 98% cache hits, same output. DeepSeek off-peak drops to $0.095 while Naive rises to $0.142, because DeepSeek's cheaper cache reads now dominate the bill. On input alone the crossover is roughly an 88% cache-hit rate. Below it, Naive's input costs less than DeepSeek off-peak. Above it, DeepSeek wins, although Naive's output stays $0.20 per million cheaper.

DeepSeek's peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, which is 9 PM-midnight and 2-6 AM Eastern. A US creator working daytime hours pays the off-peak rate. An agent fleet running overnight on a schedule may not.

3D stepped platforms showing per-task costs of $0.081, $0.089, $0.156 and $0.178
One 3M-token agent task at 90% cache hits: Naive, DeepSeek off-peak, GLM-5.3-Flash, DeepSeek peak.

Can You Run It Yourself?

Not on a workstation. The model card requires FP8-capable NVIDIA GPUs and puts the FP8 weights at about 315 GB before any KV cache. Here is the arithmetic on common rental configurations:

  • 4x H100 80GB (320 GB): the weights alone nearly fill it, so there is no room for context.
  • 4x H200 141GB (564 GB): fits, with about 249 GB left for KV cache. DeepSeek V4.1 Flash's 510.3 GB leaves only about 54 GB on the same box.
  • 8x H100 80GB (640 GB): comfortable for either model.

The only documented path today is Hugging Face Transformers 5.17.0 or later with trust_remote_code, which is fine for testing but slow for serving. NaiveAI reports 2,122 tokens per second single-stream on eight GPUs with NaiveRT, and a 3.4 ms speculative round against 12.3 ms in SGLang. NaiveRT is not public yet. The research blog says the code is planned for GitHub by 12 October, so until then any self-hosted speed you measure will not match the launch numbers.

3D slabs comparing 315 GB and 510 GB model weight sizes
FP8 weights: Naive-N0.5-Flash 315 GB, DeepSeek V4.1 Flash 510 GB.

The "Built With AI" Claim, Measured

NaiveAI's tagline is "Building Frontier AI with AI", and the blog gives concrete numbers for one part of that. NaiveRT was built in 6 days through 151 documented optimization trials. Of those, 63 changes were adopted, 71 failed or were rolled back and 17 were exploratory. Models wrote the kernels, profiled and validated them, and humans set the direction. One result was fusing DeepSeek Sparse Attention's 29 separate kernels in SGLang into a single mega-kernel.

RuntimeWire's report notes that the company has not said how much of the overall work AI did, or whether it cut costs, and that the speed figures are company-reported under narrow test conditions. The claim is credible for the serving stack, where 151 logged trials is a real audit trail. It says less about the model weights themselves. Treat it as a method worth watching, not a reason to switch models.

Which One to Point Your Agent At

For most builders this week, DeepSeek V4.1 Flash is still the default cheap open coding model. It has a live API, leads on matched Claude Code scores, and its cache pricing wins long sessions. Naive-N0.5-Flash earns a slot in three cases:

Step 1: Check your cache-hit rate. Most agent dashboards and API usage exports show cached against uncached input. If you run short tasks under about 88% cache hits, Naive's announced rates beat DeepSeek off-peak.

Step 2: Test it on greenfield builds. NL2Repo is the one benchmark where Naive clearly leads. If your work is spinning up new apps or sites from a spec rather than patching existing repos, that is the task to try it on first.

Step 3: Hold self-hosting until 12 October. Without NaiveRT you are serving a 315 GB model through Transformers. Wait for the runtime, then benchmark on 4x H200 against DeepSeek on the same box.

Step 4: Run both in your own harness. The 8.7-point harness spread above is larger than the gap between the models. Whatever you use daily, whether Claude Code, OpenCode or Codex, is the only score that predicts your results.

Frequently asked questions

What is Naive-N0.5-Flash?

An open-weight mixture-of-experts coding model from NaiveAI, released 27 September 2026. It has 309B total parameters, 15.5B active per token, a native 1M-token context and an MIT license, and it builds on Xiaomi's MiMo-V2.5 base model.

Is Naive-N0.5-Flash better than DeepSeek V4.1 Flash?

Not on most coding benchmarks. In Claude Code it scores 67.8 on DeepSWE v1.1 against DeepSeek's 69.8, and 86.7 on Terminal-Bench 2.1 against 88.0. It leads on NL2Repo-Bench, 71.9 against 64.0.

How much does the Naive-N0.5-Flash API cost?

NaiveAI lists $0.10 per million input tokens, $0.40 per million output and $0.01 per million cache reads. The model card says API access will be provided. We could not confirm open self-serve sign-up at publication.

What hardware do I need to run Naive-N0.5-Flash?

FP8-capable NVIDIA GPUs with room for about 315 GB of weights plus KV cache. Four H200s or eight H100s work. Four H100s do not leave room for context.

Can I use Naive-N0.5-Flash commercially?

Yes. The weights and inference code are released under the MIT license, which permits commercial use, modification and redistribution.

Why do benchmark scores for the same model differ between reports?

The agent harness changes the result. DeepSeek V4.1 Flash scores between 65.5 and 74.2 on DeepSWE v1.1 depending on the scaffold, so compare models only on scores from the same harness.