OpenAI opened its Decisions API to every developer as a public beta on 6 October 2026. It is a new POST /v1/decisions endpoint that runs gpt-6-luna and returns typed answers instead of text: a probability that a statement is true, a pick from a fixed list, or a score on a rubric. Input costs $0.10 per million tokens, and OpenAI charges nothing for output or caching. The API changelog says it turns text and images into answers "10x faster than the Responses API". OpenAI previewed it at DevDay in late September with limited access; this is the first version anyone can call.

What shipped

A request has three parts: model (only gpt-6-luna for now), input (a string, or user messages mixing text and images) and questions. Each question has a name and one of three types. A predicate returns a probability from 0 to 1. A choice returns one of your values plus a probability for every option and a confidence figure. A score returns the probability-weighted average of ordered levels, so a severity check can come back as 1.1 on a 0 to 2 scale. Several independent questions can share one input in a single call, and the API reference adds a fourth answer type, a refusal, for a question the model declines while the rest of the request still answers.

Two limits matter for media work. Images must be inline base64 data URLs: hosted image URLs and file IDs are rejected. And the guide says GA is expected "in the coming weeks", so the request shape can still change. Zero Data Retention, HIPAA use for eligible customers, and US or EU data residency are supported.

Original check: price and specs against the old route and two rivals

We set the new endpoint against the way you would have done this yesterday (Luna through the Responses API) and the two hosted decision models we have already tested, Cloudflare Clef and TypeSafe Jev. Every value comes from the vendor page linked in the last row. The cost row is our arithmetic: 1,000 calls of 500 input tokens each, which is 500,000 tokens.

OpenAI DecisionsLuna via ResponsesCloudflare ClefTypeSafe Jev
Input, per 1M tokens$0.10$0.10$0.24$0.042
Other chargesNone$0.50 per 1M output, cache writes $0.125Priced per input tokenPriced per input token
1,000 calls x 500 tokens$0.05$0.05 plus output and reasoning tokens$0.12$0.021
ImagesYes, base64 onlyYes, input onlyUp to 4, base64 onlyNo, text only
Context window1,050,000 (Luna)1,050,00065,53632K (per Cloudflare)
Answer typespredicate, choice, scoreFree text or JSON schemanoul, choice, scorenoul, choice, score
WeightsClosedClosedApache 2.0API only
SourceOpenAI guideOpenAI pricing, model pageWorkers AITypeSafe, Cloudflare blog

Three things fall out. Per input token, OpenAI sits exactly between the rivals: Jev is 2.4 times cheaper, Clef 2.4 times dearer. Against the old route the input price is identical, so the saving is everything else: Luna's model page lists medium reasoning effort as the default, and on the Responses API those reasoning tokens bill as output. Finally, the request shape is not TypeSafe's. OpenAI sends input and a questions array with names and calls a yes or no a predicate; Clef and Jev take a state and a map of questions keyed by id, with noul for yes or no. Code written for Jev, or for a local runtime such as Ollaya, needs a small adapter.

Input price per 1M tokens: Jev $0.042, OpenAI Decisions $0.10, Cloudflare Clef $0.24
Input price per million tokens: TypeSafe Jev $0.042, OpenAI Decisions $0.10, Cloudflare Clef $0.24.

What it means for your workflow

Decision models fit the dull checks that sit between generation steps. For creators that means jobs like:

  • Tagging a media library with a choice question (portrait, product, landscape, other) on every image.
  • Screening generated images with a predicate such as "Does any hand have the wrong number of fingers?" before a human looks.
  • Scoring thumbnails or drafts against a three-level rubric, then sending only the low scores back for another pass.
  • Sorting comments or support mail into queues, with an "other" option for anything that fits nowhere.

Because images must be inlined, a batch job has to base64-encode every file first, and the guide does not say how many input tokens one image costs. The fastest way to try it from a terminal is Simon Willison's llm-openai-decisions plugin, released a couple of hours after the beta opened: llm install llm-openai-decisions, then llm -m openai-decisions/gpt-6-luna with your question passed by -s.

What is not known yet

OpenAI gives the speed claim only as a ratio. There is no published latency in milliseconds, region or concurrency for it, and no independent result on the community Decision Index that we used to rescore Clef. The guide also does not say how many questions fit in one request (Clef allows 64). Regional processing premiums and long-context multipliers also apply on top of the $0.10 rate. We will update this page when the API reaches GA.