Comfy API went live on September 30, 2026 for everyone on a paid Comfy plan. It packages a ComfyUI workflow with its custom nodes, models and Python dependencies, then serves it as an autoscaling endpoint billed by the GPU second. We compared its rates with Runpod Serverless and Modal. Every one of Comfy's four GPU prices is exactly 1.30x Runpod's serverless flex price for the same card, to the cent.
We also ran the new comfy build init on a test ComfyUI v0.38.1 install. It pinned custom nodes to exact commits in 0.37 seconds. It also listed a public Hugging Face model as a file to upload, and it passed a spec that left out 6 packages our two node packs declare. Below: what the markup buys you, where the idle window costs more than the markup, and how to package a Build that actually runs.
What Comfy API launched on September 30
Comfy API is the deployment half of Comfy's new Developer Platform. The model is three objects. A Build describes the environment: ComfyUI version, custom nodes, models, and pip dependencies. A release is an immutable snapshot of that Build for one target, currently linux/nvidia. A deployment runs one release on one GPU type in one region, with a public URL and worker bounds set by --min and --max.
You can create a Build three ways, according to the Comfy API quickstart: scan a local install with comfy build init, import a Comfy Desktop snapshot, or drop a workflow JSON into the web Build Wizard. The launch post pitches a fourth: paste Comfy's prompt into a coding agent and let it run the CLI for you.
Pricing sits on the Comfy API product page: four GPU classes billed per second, plus network storage for models at $0.20 per GB per month. You need a paid plan to get in (Standard is $20 a month, or $16 billed yearly), and the launch post says usage is billed separately from it. Plan limits cap the rest: Standard and Creator allow 2 deployments with 2 workers each, Pro allows 5 deployments with 10 workers, Team allows 20 with 20.
Comfy API vs Runpod vs Modal: the GPU price table
The launch post names the competition itself: "If you already run ComfyUI on RunPod or Modal and it works, keep it." So we lined up the list prices. Runpod's are the serverless flex rates on the Runpod pricing page. Modal's are the per-second GPU rates on Modal's pricing page, multiplied out to an hour.
| GPU | Comfy API | Runpod Serverless (flex) | Modal (GPU only) | Comfy vs Runpod | Comfy vs Modal |
|---|---|---|---|---|---|
| RTX PRO 6000, 96 GB | $4.54/hr | $3.49/hr | $3.03/hr | 1.30x | 1.50x |
| H100, 80 GB | $6.23/hr | $4.79/hr | $3.95/hr | 1.30x | 1.58x |
| H200, 141 GB | $7.71/hr | $5.93/hr | $4.54/hr | 1.30x | 1.70x |
| B200, 180 GB | $11.23/hr | $8.64/hr | $6.25/hr | 1.30x | 1.80x |
| Model storage | $0.20/GB-mo | $0.07/GB-mo | $0.09/GiB-mo, first 1 TiB free | 2.9x | n/a |
Multiply each Runpod rate by 1.30 and round to the cent, and you get Comfy's price on all four rows: $3.49 becomes $4.54, $4.79 becomes $6.23, $5.93 becomes $7.71, $8.64 becomes $11.23. The ratios run from 1.2998 to 1.3009. That is a pricing formula, not a coincidence.
Comfy has not said which GPU provider runs a deployment. The comfy-cli 1.22.0 source has one clue. Its deployment status code translates worker counts "the same words on every GPU provider", and the fallback path carries the comment "RunPod counts one warm worker as both idle and ready". Read the markup as Comfy's price for the packaging layer, not as a faster GPU.
Modal needs one caveat. Its GPU rate excludes CPU and memory, which it bills separately. A container reserving 2 physical cores and 16 GiB adds about $0.22 an hour. We include that in the scenarios below. Modal's Starter plan also includes $30 of compute a month, and Runpod has no plan fee at all.

The 30-second idle window costs more than the markup
Per-second billing only helps if the worker stops soon after the job. The Comfy API deployment guide says a flex worker is billed from start, through the job, "plus a short idle window (currently 30 seconds)" before it scales down. Neither comfy deploy up nor comfy deploy scale has an option to change it. Runpod's default idle timeout is 5 seconds and configurable, per the Runpod Serverless billing docs. Modal's default is 60 seconds, adjustable from 2 seconds to 20 minutes with scaledown_window, per Modal's cold-start guide.
We modelled one month on an RTX PRO 6000 at list price. Cold-start time is excluded for all three because none publishes it, so real bills will be higher everywhere.
| Scenario (RTX PRO 6000) | Comfy API | Runpod flex | Modal (+2 cores, 16 GiB) |
|---|---|---|---|
| 1,000 jobs of 20 s, each arriving alone (default idle windows) | $63.06 | $24.24 | $72.30 |
Same, Modal scaledown_window=2 | $63.06 | $24.24 | $19.88 |
| 1,000 jobs of 20 s in one batch | $25.26 | $19.39 | $18.13 |
| One warm worker, 730 hours | $3,314.20 | $2,547.70 | $2,374.98 |
| 50 GB of models, one month | $10.00 | $3.50 | $0 (inside 1 TiB free) |
In the first row, 30 of every 50 billed seconds on Comfy API are idle: 60% of the bill. That turns a 1.30x rate gap into a 2.6x bill gap against Runpod. Batch the same jobs and the gap falls back to 1.30x. So the cheapest way to use Comfy API is to send work in bursts, or to run it on internal tools where a few seconds of queueing do not matter.
Always-on traffic works the other way. --min 1 keeps a worker warm around the clock, and at $4.54 an hour that is $3,314 a month before storage. Runpod sells discounted active workers through sales, so its row is a ceiling.

What comfy build init captured in our test
The pricing only matters if the packaging works, since that is what the markup pays for. We installed comfy-cli 1.22.0 (published on PyPI at 01:00 UTC on September 30) and built a test install. It had ComfyUI v0.38.1, two popular public node packs (KJNodes and VideoHelperSuite) and one private node folder with no git remote. For models it had the public TAESD decoder from Hugging Face plus a 3 MB private file named like a house-style LoRA. Then we ran the scan against a Python venv, as --python asks.
What it got right:
- Exact node commits. Both public packs were recorded with their repository URL and full 40-character commit SHA, not a branch name. The Build installs the code you tested.
- Private nodes travel. The folder with no remote was recorded as
source: localwith a SHA-256 digest of its 453 bytes, and the dry-run push counted it as an upload. - Models by hash. Both model files were recorded with SHA-256 and byte size, so a changed checkpoint shows up as a changed Build.
- Speed. The whole scan took 0.37 seconds.
What needs a second look:
- Public models get uploaded. The TAESD decoder is on Hugging Face, but the offline scan marked it
source: local, the same as the private file. The dry run planned 3 uploads totalling 7.9 MB. For a 20 GB checkpoint that means a long upload, and Comfy's ownvalidate --remotelookup of public copies needs a signed-in account. - Pip pins come from your venv, not your nodes. The spec recorded 8 packages, exactly what we had installed. The two node packs' requirements files name 6 packages that venv lacked (opencv-python, opencv-python-headless, matplotlib, color-matcher, mss and imageio-ffmpeg), and
comfy build validatepassed anyway. Once we installed those requirements and rescanned, the spec had 25 pins. It now included bothopencv-pythonandopencv-python-headless5.0.0.93, because KJNodes asks for both. - Torch is whatever you have. Our venv had no torch, and the spec recorded
torch: nullwithout a warning. The captured file's own header says pins "from where the workflow ran locally" must be retargeted for a different OS or GPU.
We did not deploy. comfy build push without an account stops with build_not_signed_in, and a real push needs a paid plan. Whether the remote builder also installs each pack's requirements.txt is not documented, so treat the spec as the source of truth until your first release log proves otherwise.

How to package a ComfyUI workflow for Comfy API
Step 1: scan the environment the workflow really runs in. Point --python at the ComfyUI venv that has torch and every node pack's requirements installed, not a fresh one. The spec records the output of pip freeze.
pip install -U comfy-cli
cd ComfyUI
comfy build init . --name product-shots --python ./venv/bin/pythonStep 2: link public models instead of uploading them. Open comfy-build.yaml. For every model that has a public https URL, delete localPath and source: local and add sourceUri. In our test that cut the dry-run upload from 3 items and 7.9 MB to 2 items and 3.0 MB, and validation still passed.
- filename: taesd_decoder.safetensors
sourceUri: https://huggingface.co/madebyollin/taesd/resolve/main/taesd_decoder.safetensors
sha256: f0fb51dd10d41c26612c070fa0b52ea0215a5ff90792134b4971109dd713c019
type: vae_approxStep 3: dry-run before you spend. comfy build validate . checks the spec offline, and comfy build push . --dry-run lists every upload and its size without sending anything. Read the pip list for packages your nodes need but the spec lacks.
Step 4: release, then read the build log. comfy build push, then comfy build release create --target linux/nvidia --watch. The release log is the first place a missing dependency shows up.
Step 5: deploy at zero and batch your calls. comfy deploy up --gpu <gpu> --region <region> --min 0 --max 2 scales to nothing when idle. Queue jobs so a warm worker takes several in a row, since every scale-down costs 30 seconds of GPU time. On an RTX PRO 6000 that is about 3.8 cents each time.
Step 6: delete what you are not using. Per the deployment guide, staged model storage is billed while any deployment of the Build exists in that region, including a paused one. Builds that were never deployed are free to keep.
Comfy API or Comfy Cloud API: which one you are buying
Comfy now sells two things with "API" in the name, and the Comfy pricing page describes both. The older Comfy Cloud API runs workflows on Comfy's shared RTX 6000 Pro machines, draws from your plan's credit pool, and allows 1, 3, 5 or 25 concurrent jobs on Standard, Creator, Pro and Team. It runs Comfy's managed environment with 900+ preinstalled models (your own models and LoRAs import on Creator and above), and runs cap at 30 minutes (1 hour on Pro).
The new Comfy API runs your own Build, custom nodes included, on a GPU you choose, billed per worker-second. Comfy prices it in the same credits, at 957.94 credits per RTX PRO 6000 hour, or about 211 credits per dollar. If a workspace runs out of credits, its deployments are stopped automatically. If your workflow runs on the stock Cloud nodes, the Cloud API is simpler and needs no Build. If it depends on custom nodes or pinned versions you need to control yourself, Comfy API is the path. Comfy's Router and Partner Node markups are a third bill again.
Who should use Comfy API, and who should not
Use it if the ComfyUI environment is the thing you keep fixing. The immutable Build plus pinned commits plus one CLI is real work saved. A small studio handing a tuned product-photo workflow to clients gets a versioned endpoint without writing a Dockerfile, and Silverside AI is the launch reference for that pattern. Budget 1.30x Runpod list plus 30 seconds per scale-down.
Skip it if you already run a ComfyUI worker on Runpod that builds cleanly. You would pay 30% more on compute and 2.9x on storage for packaging you have already done. Skip it too for spiky, one-request-at-a-time public traffic, where the fixed idle window can put your bill at 2.6x Runpod's. For always-on traffic, a warm Runpod or Modal worker is about $770 to $940 a month cheaper at list.
For builders comparing sandboxes and GPUs for agents rather than image pipelines, our Docker Cloud Sandboxes vs E2B, Vercel and Modal cost test covers the same per-second trade-offs.
Frequently asked questions
How much does Comfy API cost per hour?
$4.54 on an RTX PRO 6000, $6.23 on an H100 SXM, $7.71 on an H200 SXM and $11.23 on a B200, billed per second, plus $0.20 per GB-month of model storage. You also need a paid Comfy plan, from $20 a month.
Is Comfy API cheaper than Runpod?
No. At list price every Comfy API GPU rate is 1.30x Runpod Serverless flex for the same card, and storage is 2.9x. With traffic that arrives one request at a time, Comfy's fixed 30-second idle window can make the bill about 2.6x Runpod's.
Can I use custom nodes and my own LoRAs on Comfy API?
Yes. comfy build init records public node packs by repository and commit, uploads private node folders by hash, and packages local model files. That control is the main difference from the Comfy Cloud API, which runs Comfy's managed environment.
Does comfy build init capture every Python dependency my workflow needs?
Only the ones installed in the Python environment you point it at. In our test it recorded 8 pins from a venv missing the node packs' requirements, and validation still passed. Scan the venv your workflow actually runs in.
Does Comfy API scale to zero?
Yes, with --min 0. Each flex worker stays billed for about 30 seconds after its last job, and the first request after idle waits for a cold start.
What is the difference between Comfy API and the Comfy Cloud API?
The Cloud API runs workflows on Comfy's shared machines from your plan credits, with 1 to 25 concurrent jobs by plan. Comfy API runs your own packaged environment on a GPU you choose, as a dedicated autoscaling endpoint billed per worker-second.