OpenAI deprecated all four of its text-to-speech models on 1 October 2026: tts-1, tts-1-hd and both gpt-4o-mini-tts snapshots stop working on January 6, 2027. The named replacement, gpt-realtime-2.1-mini, is not a drop-in. It does not serve the /v1/audio/speech endpoint those models use, it cannot return MP3, and it has no fable, onyx or nova voice. Changing one model string fixes nothing.
We read OpenAI's model pages, API reference and pricing, and worked out what narration costs after the switch. We also checked the source of nine popular open-source creator tools, seven of which default to a retiring model. Then we tested the cheapest way out: OpenAI's own sample call, unchanged except for one line, against a free local server running on a CPU. No OpenAI call was made, so nothing here was paid for.
What OpenAI deprecated on October 1
The notice sits on OpenAI's deprecations page under "2026-10-01: Text-to-speech models". It gives "at least three months' notice" and one instruction: "Migrate to gpt-realtime-2.1-mini before the shutdown date." That leaves 96 days from today.
| Model | Shutdown | Price today | Replacement |
|---|---|---|---|
tts-1 | Jan 6, 2027 | $15 per 1M characters | gpt-realtime-2.1-mini |
tts-1-hd | Jan 6, 2027 | $30 per 1M characters | gpt-realtime-2.1-mini |
gpt-4o-mini-tts-2025-03-20 | Jan 6, 2027 | $0.60 text in, $12 audio out per 1M tokens | gpt-realtime-2.1-mini |
gpt-4o-mini-tts-2025-12-15 | Jan 6, 2027 | $0.60 text in, $12 audio out per 1M tokens | gpt-realtime-2.1-mini |
The detail that matters most is in the speech endpoint's API reference. Its model field accepts exactly four values: tts-1, tts-1-hd, gpt-4o-mini-tts and gpt-4o-mini-tts-2025-12-15. All four are on the shutdown list. Unless OpenAI adds a model before January, /v1/audio/speech has nothing left to run. As of today the text-to-speech guide still recommends gpt-4o-mini-tts and carries no deprecation notice at all.

The replacement speaks a different API
The gpt-realtime-2.1-mini model page lists one supported endpoint: v1/realtime, reached over WebRTC, WebSocket or SIP. Speech generation, Chat Completions, Responses and Batch are all marked "Not supported". It is a conversational voice model with reasoning, built for live agents. OpenAI launched it in July as the low-cost tier of gpt-realtime-2.1, not as a narration engine.
Today: /v1/audio/speech | After: gpt-realtime-2.1-mini | |
|---|---|---|
| How you call it | One HTTP POST, audio file back | Open a socket session, send events, collect audio deltas |
| Output formats | MP3, Opus, AAC, FLAC, WAV, PCM | PCM at 24 kHz, G.711 mu-law or a-law |
| Built-in voices | 13 | 10 |
| Longest single request | 4,096 characters (tts-1), 2,000 tokens (gpt-4o-mini-tts) | 32,000 output tokens, about 26.7 minutes of audio |
| Batch API | No | No |
| Reads your text as written | Yes, that is the whole job | Only if a prompt tells it to |
That last row is the one that changes narration work. OpenAI's Realtime conversations guide shows how to make the model read fixed text: send a response.create event with an empty input array and instructions that begin "Say exactly the following:". It is a prompt to a reasoning model, not a guarantee. For a script you have already approved, every take now needs checking against the text.
Formats change too. The Realtime API returns raw PCM, so anything that expected an MP3 file now needs an encoding step, usually ffmpeg. The longer limit is the one real gain: a single response can carry about 26 minutes of audio, where tts-1 stops at 4,096 characters, roughly 679 words of ordinary prose.
Three voices have no successor
The speech endpoint offers 13 built-in voices. The Realtime session reference lists 10: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin and cedar. fable, onyx and nova are missing.
Those three date back to the original tts-1 launch, and they are the ones a lot of faceless YouTube channels, audiobook pipelines and podcast intros were built on. After January 6 there is no OpenAI model that speaks in them. If your channel's sound is onyx, you are choosing a new voice either way, so pick it now and re-record the intro on your schedule rather than OpenAI's.
Custom voices are not part of this. OpenAI's custom voices guide says an approved custom voice works with the speech endpoint, the Realtime API and Chat Completions audio output, so an eligible account keeps its voice ID.

What narration costs after the switch
tts-1 bills by input character. gpt-realtime-2.1-mini bills by token: $0.60 per million text tokens in and $20 per million audio tokens out, according to OpenAI's pricing page. OpenAI's Realtime cost guide says assistant audio counts as "1 token per 50ms", which is 1,200 tokens, or $0.024, per minute of speech.
So the comparison depends on how fast the voice talks. We measured 6.03 characters and 1.23 tokens per word on 7,274 words of English sentences, then priced a 1,000-word script at three speaking rates.
| 1,000-word script | Audio length | tts-1 | tts-1-hd | gpt-realtime-2.1-mini | Change vs tts-1 |
|---|---|---|---|---|---|
| 130 words a minute | 7.7 min | $0.090 | $0.181 | $0.185 | +105% |
| 150 words a minute | 6.7 min | $0.090 | $0.181 | $0.161 | +78% |
| 170 words a minute | 5.9 min | $0.090 | $0.181 | $0.142 | +57% |
Moving from tts-1 costs 57% to 105% more per script, more for slow, deliberate reads. tts-1-hd users come out 11% cheaper at 150 words a minute. These are list prices for audio that comes out right the first time; any retake to fix a paraphrased line is billed again. gpt-4o-mini-tts users face a similar rise, since its $12 audio rate becomes $20, though OpenAI's current pages do not state how many tokens a minute of its audio uses.
There is one other OpenAI route the notice does not mention. gpt-audio-mini runs on plain HTTP Chat Completions at the same $20 audio rate and returns the audio inside the JSON response. It is also a conversational model answering a prompt, so the same check-every-take caveat applies.

We checked 9 open-source creator tools
We read the current source of nine widely used tools that can speak through OpenAI. All nine call the /audio/speech route that loses its models in January. Seven of the nine default to a model on the shutdown list.
| Tool (GitHub stars) | What the code does | Way out without a rewrite |
|---|---|---|
| n8n (206,506) | OpenAI node "Generate audio" defaults to tts-1; the only options are tts-1 and tts-1-hd | OpenAI credential has a Base URL field |
| Open WebUI (153,806) | AUDIO_TTS_MODEL defaults to tts-1 | AUDIO_TTS_OPENAI_API_BASE_URL |
| MoneyPrinterTurbo (128,059) | Uses /audio/speech for self-hosted voices; no OpenAI default | Already points at local servers |
| NextChat (88,819) | Default tts-1; model list is tts-1 and tts-1-hd | Custom OpenAI endpoint |
| LobeChat (82,946) | Default ttsModel: 'tts-1'; choices are all three retiring models | Proxy URL for the OpenAI provider |
| AnythingLLM (66,668) | Built-in OpenAI voice provider hardcodes model: "tts-1", with no setting to change it | Switch to its "OpenAI Compatible" provider |
| LibreChat (45,203) | Model comes from your config; URL defaults to OpenAI's speech endpoint | url in the speech.tts config |
| SillyTavern (34,024) | OpenAI TTS defaults to tts-1; every choice is a retiring model or an older tts-1 snapshot | "OpenAI Compatible" TTS provider |
| openai_tts for Home Assistant (219) | Falls back to tts-1; new setups default to gpt-4o-mini-tts | Custom URL in the integration |
None of them can reach gpt-realtime-2.1-mini through these settings, because it lives on a different endpoint. Until each project ships a Realtime client, the practical exits are the base URL fields in the last column. Commit hashes and the matching lines are in our audit files. Stars and code were read on 2 October.

Tested: keep your code, swap the server
Every tool above speaks one simple protocol, so the cheapest migration is to keep that protocol and change who answers it. Kokoro-FastAPI (Apache 2.0, 5,504 stars) wraps the 82-million-parameter Kokoro model, also Apache 2.0, behind an OpenAI-compatible /v1/audio/speech endpoint. Kokoro ranked first in English on the open leaderboard we covered in our Open TTS Leaderboard licence test.
We ran version 0.9.1-rc1 on CPU and sent it OpenAI's own sample call from the text-to-speech guide, using the official openai Python SDK 3.24.0. The only edit was the client line:
client = OpenAI(base_url="http://127.0.0.1:8880/v1", api_key="not-needed")
response = client.audio.speech.create(model="tts-1", voice="onyx", input=text)- Models:
tts-1,tts-1-hdandgpt-4o-mini-ttswere accepted. The dated namegpt-4o-mini-tts-2025-12-15returned a 400 error, so use the short alias. - Formats: all six worked. MP3, AAC, FLAC, WAV and PCM came back at 24 kHz, and Opus at 48 kHz.
- Voices: 9 of OpenAI's 13 names worked, including
fable,onyxandnova.ballad,verse,marinandcedarreturned 400 because they are not in the server's mapping file. They map to Kokoro voices, not to OpenAI's sound:onyxbecomes British malebm_george,alloybecomesam_adam. - Instructions: the
instructionsfield was accepted without error and ignored; the server code never reads it. - Length: a 1,006-word script of 6,112 characters, past OpenAI's 4,096-character cap, went through in one request. It came back as 6 minutes 47 seconds of MP3, spoken at 148 words a minute, close to the middle row of the cost table.
That render took 48.8 seconds, 8.3 times faster than real time, while an unrelated job was using all six cores of the test machine, so a quiet machine will do better. On the same CPU without contention in September, Kokoro ran 10.5 times faster than real time with a median 366 ms per sentence in our leaderboard test. The cost per minute is your electricity.
How to migrate before January 6
Step 1: Find every call
Search your code, automation exports and tool settings for the four model names and the endpoint: grep -rnE "tts-1|gpt-4o-mini-tts|audio/speech" . In n8n, export your workflows to JSON first and search those. Check environment files too; Open WebUI keeps its choice in AUDIO_TTS_MODEL.
Step 2: Check your voice
If any call uses fable, onyx or nova, that voice ends with the model. Audition the replacement now on a real script, whichever path you take.
Step 3: Pick a path
Three options keep narration running. A Realtime rewrite keeps OpenAI's ten voices and quality but means socket code, PCM output and checking each take. Another hosted service with an OpenAI-compatible speech endpoint keeps your code. A local Kokoro-FastAPI server keeps your code and costs nothing per minute.
Step 4: Run the local server
The project's quickest start is one command: docker run -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-cpu:latest. Check http://localhost:8880/health returns "healthy".
Step 5: Point your tool at it
Set the base URL to http://localhost:8880/v1 in the field from the table above, keep tts-1 as the model name, and send a short test line. In n8n, create a separate OpenAI credential for this, because the Base URL applies to every call that uses the credential. In AnythingLLM, move from the "OpenAI" voice provider to "OpenAI Compatible".
Step 6: Fix the voice names
Either use Kokoro's own voice names directly, or edit api/src/core/openai_mappings.json so the OpenAI names your tools send point to the Kokoro voices you picked in Step 2.
Troubleshooting
The server returns 400 "Voice not found"
The name is not in the mapping file. ballad, verse, marin and cedar fail out of the box.
The server returns 400 "Unsupported model"
You are sending a dated snapshot. Use tts-1 or gpt-4o-mini-tts.
The tone instructions do nothing
Kokoro does not support them. Pick a voice that already has the delivery you want.
Frequently asked questions
When does OpenAI's tts-1 stop working?
On January 6, 2027. The same date applies to tts-1-hd and both gpt-4o-mini-tts snapshots, per OpenAI's deprecation notice of October 1, 2026.
Can I just change the model name to gpt-realtime-2.1-mini?
No. That model does not support the /v1/audio/speech endpoint. It only runs on the Realtime API, which uses a socket session and returns PCM audio rather than an MP3 file.
Is gpt-realtime-2.1-mini more expensive than tts-1?
Yes, for most narration. At OpenAI's list prices a 1,000-word script costs about $0.090 on tts-1 and $0.142 to $0.185 on gpt-realtime-2.1-mini, depending on speaking speed. It is slightly cheaper than tts-1-hd at normal speed.
What happens to the onyx, nova and fable voices?
They are not offered on the Realtime API, so they end with the speech models on January 6. Custom voices created from audio samples carry over.
Will n8n, Open WebUI and other tools update?
That is up to each project. As of 2 October, seven of the nine tools we checked still default to a retiring model, and none had a Realtime voice path for narration.
Is a local TTS server good enough for paid work?
Kokoro and Kokoro-FastAPI are both Apache 2.0, which allows commercial use. Quality is a judgment call on your own scripts; listen to a full take before you switch a channel over.