Together AI launched Together Link on October 5, 2026, a one-command installer that points Claude Code, Codex, OpenCode, Pi Code, Claude Desktop and ChatGPT Desktop at open models hosted on Together: Kimi K3, GLM 5.3, GLM 5.3 Flash and DeepSeek V4.1 Flash. Together says it cuts coding-agent spend "by over 50%" against Claude Opus 5.5. We priced 13 real Opus 5.5 sessions from our own Claude Code automation (597 API calls) at each model's published rates. The 50% holds on the two Flash models (92% and 96% cheaper). Kimi K3, which Together Link puts in Claude Code's "Opus" slot, comes out only 6% cheaper, and on our longest session it costs more than Opus 5.5.
The reason is the part of the bill most comparisons skip. In an agent session, 98% of input tokens are cache reads, and Anthropic charges $0.20 per million for an Opus 5.5 cache read while Together charges $0.30 for Kimi K3 and $0.26 for GLM 5.3. Below: what Together shipped, the session math, what the Claude Code model menu actually runs, and how to check your own savings before you switch a team over.
Background: what Together Link does
Together Link is a launcher, not a new agent. You install it with curl -fsSL https://link.together.ai/install | bash, save a Together API key with togetherlink configure, then start your usual tool through it: togetherlink claude (or the shortcut tclaude), togetherlink codex, togetherlink opencode or togetherlink pi. The Together Link docs label the product beta and say commands, routing and the model list may change.

For terminal agents, Together Link passes a temporary configuration for that launch only, and your normal settings files are left alone. Claude Desktop and ChatGPT Desktop get a separate Together Link profile you can switch off with togetherlink claude-desktop off or togetherlink chatgpt off. Billing runs on your existing Together key, pay as you go or credit packs, with no contract. Each session prints a cost receipt when it exits, and in Claude Code the status line shows the session's spend next to what the same tokens would cost on Opus.
Requirements are modest: macOS or Linux with Bash and curl, the agent already installed, OpenCode 2 (version 1 is refused), and Pi Code 0.80.8 or newer. The installer also installs Bun if it is missing. The support repository on GitHub is MIT licensed, but it holds the README and issue tracker, not the CLI source; the CLI ships as a compiled JavaScript bundle (v0.9.75 at the time of writing).
Deep Analysis
We read the launch post, the docs, the README and the CLI bundle itself, then ran the numbers on real sessions. Four findings matter more than the headline.
The cache bill is the real bill
A coding agent resends its whole conversation on every turn: system prompt, tool definitions, every file it read, every command output. Prompt caching turns most of that into cheap cache reads. In our 13 Opus 5.5 sessions (October 4 and 5, content and operations work in Claude Code, 10 to 100 API calls each), the averages were about 131,000 input tokens and 621 output tokens per call. Of 77.96 million input tokens, 76.37 million were cache reads (98.0%), 1.59 million were cache writes and 1,202 were uncached. Output was 0.47% of all tokens.
Here is how those prices compare per million tokens, from Together's pricing page and Anthropic's pricing page:
| Model | Input | Cache read | Output |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | $4.00 | $0.20 | $20.00 |
| Kimi K3 (Together) | $3.00 | $0.30 | $15.00 |
| GLM 5.3 (Together) | $1.40 | $0.26 | $4.40 |
| GLM 5.3 Flash (Together) | $0.15 | $0.03 | $0.50 |
| DeepSeek V4.1 Flash (Together) | $0.30 | $0.006 | $1.20 |
Opus 5.5's cache read is 5% of its input price; Kimi K3's is 10% and GLM 5.3's is 19%. Anthropic does charge for cache writes ($8 per million for the one-hour cache Claude Code used in every one of our sessions), and Together does not charge a write premium. That is the trade: Together wins on writes and output, Anthropic wins on reads, and reads are 98% of the volume.
What 13 real sessions would cost
We priced the exact token counts from the session logs at each rate. Together's own CLI bundle carries a slightly lower Kimi K3 rate ($2.70 input, $0.27 cache read, $13.50 output) than its public pricing page, so we show both.

| Model | 13 sessions | Per session | vs Opus 5.5 |
|---|---|---|---|
| Claude Opus 5.5 | $35.41 | $2.72 | baseline |
| Kimi K3 (pricing page) | $33.24 | $2.56 | -6% |
| Kimi K3 (CLI table) | $29.92 | $2.30 | -16% |
| GLM 5.3 | $23.71 | $1.82 | -33% |
| GLM 5.3 Flash | $2.72 | $0.21 | -92% |
| DeepSeek V4.1 Flash | $1.38 | $0.11 | -96% |
Session length changes the answer. On short sessions, where cache writes are a large share of the bill, Kimi K3 wins: our 10-call session was $0.59 on Opus 5.5 and $0.37 on Kimi K3. On the longest one, 100 calls and 17.8 million cache reads, Opus 5.5 was $6.68 and Kimi K3 was $6.94. With these prices, Kimi K3 is cheaper only while a session's cache reads stay below about 50 times its cache writes plus output. Long, tool-heavy agent sessions cross that line.
Three caveats keep this honest. Different models use different tokenizers, so the same text is not the same token count. An open model may take more or fewer turns to finish the same task. And the comparison assumes Together's cache hits as often as Anthropic's, which Together does not promise: its prompt caching docs call caching "best effort" and say entries may be evicted sooner under heavy load. If every request missed the cache, the same tokens would cost $239.45 on Kimi K3, almost seven times the cached Opus 5.5 bill.
Three sources, three routing stories
"Auto" is the default model, and how it routes decides whether the cache survives. The launch post says the router "reads each session's first task" and that "routing happens once per session, so prompt caching keeps working." The docs say that "for each request," the gateway classifies the prompt and routes it. The README's model table says Auto "routes per request," and the CLI's own setup screen says an Anthropic key "lets Auto send the hardest turns to Claude."
That difference is not a detail. Together's caching docs are explicit that "a cache entry only serves requests to the same model." If Auto switches models mid-session, the new model starts with a cold cache and the whole conversation is billed at full input price on that turn. Per-session routing keeps the savings; per-request routing can erase them on a long session. Until Together reconciles the three descriptions, measure it yourself with togetherlink usage (see the steps below).
Auto's lineup also depends on your keys. With an Anthropic key saved, Claude Code and Claude Desktop sessions route between GLM 5.3 and Opus 5.5, and Opus turns are billed to your Anthropic account. Without one, the launch post says Auto routes between GLM 5.3 and GLM 5.3 Flash. On our mix, if Auto routes whole sessions, it has to send roughly 3 in 10 of them to GLM 5.3 Flash to clear the 50% line.
What the Claude Code model menu actually runs
Inside Claude Code, Together Link relabels the familiar tiers. In the shipped CLI (v0.9.75) and the README, picking Opus runs Kimi K3, Fable runs GLM 5.3, Sonnet runs DeepSeek V4.1 Flash and Haiku runs MiniMax M3. The docs page lists a different mapping for the two lower tiers (Sonnet on GLM 5.3 Flash, Haiku on DeepSeek V4.1 Flash). The banner at the start of each session names the route, so check it rather than trusting the menu label.
This matters for more than price. Claude Code uses its smaller tiers for background jobs. In August, the Together Link team found that Claude Code's Auto-mode Bash safety classifier, which runs on the Sonnet tier, "intermittently timed out or returned invalid output" on MiniMax M3; pull request #21 moved those classifier calls to DeepSeek V4 Flash. That is the kind of failure to watch for: not the main model getting a task wrong, but a helper model quietly behaving differently.
What carries over and what does not
Most of your setup survives the swap. In OpenCode, your MCP servers, skills, plugins and agents still load. In Pi Code, extensions, MCP servers, skills and prompt templates load as usual. Codex's native web search works. Pin a model for a whole launch with togetherlink --main zai-org/GLM-5.3 claude; the flag must come before the tool name, and the shortcuts such as tclaude cannot take it.

A few things change. ChatGPT Desktop's native web search is disabled in the Together Link profile. OpenCode's web search uses whatever integration you connect inside OpenCode, not the gateway. Claude Code, Claude Desktop and ChatGPT Desktop sessions gain an image generation skill that bills separately from tokens. Headless Claude Code runs hang unless you append < /dev/null. And every request, including the Opus turns paid with your own Anthropic key, passes through Together's hosted gateway, so check that against your code-handling rules before you route client work through it.
One earlier bug is worth knowing about if you tried a pre-launch build. Issue #23, filed September 16, showed the wrapper let Bun load the launch folder's .env file into the Claude Code session, including any ANTHROPIC_API_KEY in it. The current installer runs Bun with --no-env-file. If you installed before late September, run togetherlink update.
How to check your own savings in one afternoon
- Measure your baseline. Claude Code writes token usage for every call into its session logs under
~/.claude/projects/. Total the input, cache read, cache write and output tokens from a typical week. If cache reads are above 95% of input, as ours were, the cache price decides the winner. - Install and pin one model. Run the installer, then
togetherlink --main zai-org/GLM-5.3-Flash claudeon a real but low-risk task. Start with a Flash model, because that is where the savings are large enough to survive tokenizer and turn-count differences. - Read the receipt. Each session prints a cost summary on exit. Run
togetherlink usage --last 7dafter a few sessions and compare against the same kind of work on Opus 5.5. - Test Auto separately. Run one long session on Auto and one pinned to GLM 5.3. If Auto costs more per call late in the session, it is switching models and losing the cache.
- Check the helpers. If you use Claude Code's Auto permission mode, watch for unexpected prompts or blocked commands; those come from the classifier running on an open model.
Impact on Creators
If you build with Claude Code for side projects, client sites or creative tools and your bill is mostly short sessions, Together Link with GLM 5.3 Flash or DeepSeek V4.1 Flash can cut it by roughly 90% on token prices alone. Our earlier GLM-5.3-Flash coding comparison and the DeepSeek V4.1 Flash test in Claude Code show what those models handle well and where they slip. Use them for scaffolding, refactors, scripts and boilerplate, and keep the hard debugging on a frontier model.
If you run long agent sessions, the kind that read a whole repository and loop for an hour, do not assume the "Opus" slot is the cheaper Opus. Kimi K3 is a strong model, covered in our Kimi K3 launch analysis, but at Together's list prices it costs about the same as Opus 5.5 once caching is counted. The Opus 5.5 price cut already did much of what you would be switching for.
Key Takeaways
1. Together Link swaps the model behind Claude Code, Codex, OpenCode and Pi with one command and leaves your normal configuration untouched.
2. On 13 real Opus 5.5 sessions, GLM 5.3 Flash was 92% cheaper and DeepSeek V4.1 Flash 96% cheaper; Kimi K3 was 6% cheaper at list price and GLM 5.3 33% cheaper.
3. Cache reads were 98% of input, and Opus 5.5's $0.20 cache read undercuts Kimi K3 ($0.30) and GLM 5.3 ($0.26), so long sessions favour Opus.
4. Together's launch post, docs and README disagree on whether Auto routes per session or per request; per-request routing would reset the cache on every switch.
5. The Claude Code menu labels are relabelled: "Opus" runs Kimi K3 in the shipped CLI, and helper jobs such as the Auto-mode safety classifier run on open models too.
What to Watch
The first thing to watch is a clear statement from Together on how Auto routes, ideally with cache-hit rates in the session receipt. If routing is per session and the receipt shows cache reads at the cached price, the Flash tiers are a straightforward saving for everyday coding. The second is pricing: Kimi K3 already appears at $2.70 in the CLI's own table against $3.00 on the public page, and a cache-read price below Opus 5.5's $0.20 would change the long-session math. Third, the model list is marked as subject to change during the beta; MiniMax M3 and Qwen 3.8 are in the CLI catalogue but not the docs table, so expect the lineup to move. Run your own receipts for a week before you move a team.
Frequently asked questions
What is Together Link?
Together Link is a launcher from Together AI, released October 5, 2026, that runs Claude Code, Codex, OpenCode, Pi Code, Claude Desktop and ChatGPT Desktop against open models hosted on Together, billed to your Together API key. It is in beta.
Which models does Together Link use?
The docs list Auto, Kimi K3, GLM 5.3, GLM 5.3 Flash and DeepSeek V4.1 Flash, each with a 1 million token context. The shipped CLI also includes MiniMax M3 and Qwen 3.8. Run togetherlink models for the live list and prices.
Does Together Link really save 50% on Claude Code?
On our 13 real sessions, only the Flash models cleared 50%: GLM 5.3 Flash saved 92% and DeepSeek V4.1 Flash 96%. Kimi K3 saved 6% at list price and GLM 5.3 saved 33%, because Opus 5.5's cache reads are cheaper than theirs.
Does Together Link change my Claude Code settings?
No. For terminal agents it passes a temporary configuration for that launch only, and your normal files are unchanged. To go back, launch your tool directly. Desktop apps use a separate profile you can turn off or reset.
Do I need an Anthropic API key?
No. Without one, every request is served by Together models. With one saved, Auto can send the hardest requests in Claude Code and Claude Desktop to Claude Opus, billed to your Anthropic account.
Do MCP servers and skills still work?
In OpenCode and Pi Code, the docs say your MCP servers, skills, plugins and extensions load as usual. How well each open model uses those tools is a separate question, so test your most important tools on a pinned model first.