Wispr's research arm published Canto on 17 September 2026, a speech model for real-time dictation, and led with one number: 3.4% word error rate on an evaluation of real dictations, the lowest of the six models it tested. The comparison charts on that page carry a second number the announcement never mentions. On short dictations, the one and two word utterances people fire off all day, Canto scores 21.4%, and so do two of the models it beat.
That is the same model, on the same evaluation, 6.3 times worse. It is also a three way tie, which means the thing a buyer is choosing between makes no measurable difference in that condition at all. The numbers below were read off the four charts published with the launch by the Wispr Advanced Interfaces Lab, because they exist only as images in that post. The arithmetic on top of them is ours.
What Wispr shipped
Canto is a dictation model, not an API product. It runs inside Wispr Flow, the dictation app the company sells, and the launch post publishes no latency figure, no token price and no endpoint. The only price attached to it is the app's: Wispr Flow's plans run from a free tier capped at 2,000 words a week on desktop and 1,000 on mobile, to Pro at $15 per user per month ($12 billed annually), Growth at $33 ($26 annually), and custom Enterprise pricing. If you want Canto, you buy the app.
The lab that built it was introduced on 24 July 2026 and is led by Ariya Rastrow, a founding member of Amazon's Alexa team. Canto is described as its first model, with a successor already training at more than ten times the scale.
The numbers, read off the charts
Wispr ran two evaluations of its own. The first is 10 hours of English Wispr Flow dictations from more than 2,300 speakers, randomly sampled. The second is a 3 hour challenge set deliberately assembled from the conditions that break dictation: nearby speech, music, traffic, wind, low recording volume, whispered or far field speech, and very short utterances.
| Model | Real-world dictation (10h) | Challenge set (3h) |
|---|---|---|
| Wispr Canto | 3.4% | 9.1% |
| Google Gemini 3.1 Pro | 3.6% | 8.2% |
| OpenAI GPT Transcribe | 4.0% | 11.3% |
| AssemblyAI Universal-3.5 Pro | 4.2% | 13.9% |
| Google Gemini 3.5 Transcribe | 4.5% | 12.2% |
| Deepgram Nova-3 | 7.4% | 23.5% |
Two things happen between those columns. The winner changes: Canto leads the random sample by 0.2 percentage points over Gemini 3.1 Pro and loses the hard one by 0.9, which is 11% relative. Wispr says so itself, and adds the fair qualifier that Gemini 3.1 Pro is a frontier-size multimodal model not suited to real-time low-latency use. The second change is the one worth acting on: the spread between best and worst widens from 2.18x to 2.87x, and the top four models stop being interchangeable. On the random sample they sit inside 0.8 percentage points of each other. On the hard set they span 5.7.

Where the models actually separate
Wispr broke the challenge set into three conditions. This table is where a purchasing decision actually lives, and it is the one that appears nowhere in the announcement text.
| Condition | Canto | Gemini 3.1 Pro | Gemini 3.5 Transcribe | GPT Transcribe | AssemblyAI | Deepgram Nova-3 |
|---|---|---|---|---|---|---|
| All challenging audio | 9.1% | 8.2% | 12.2% | 11.3% | 13.9% | 23.5% |
| Noisy | 7.6% | 6.1% | 10.0% | 8.9% | 10.4% | 22.2% |
| Low volume | 10.1% | 10.1% | 14.1% | 13.2% | 18.3% | 25.0% |
| Short dictation | 21.4% | 21.4% | 25.0% | 33.6% | 21.4% | 41.4% |
Noise is where the money is. Best to worst on noisy audio is 6.1% against 22.2%, a factor of 3.64, the widest gap in the entire release. If you dictate on a commute, in a shared office or anywhere with a second voice in the room, that row decides your experience and the headline row does not. Low volume, meaning whispered or far-field speech, spans 2.48x and produces an exact tie at the top between Canto and Gemini 3.1 Pro.
Short dictation is where everyone flattens out
The last row behaves differently from the other three. Wispr defines short dictations as one to two word samples with little surrounding context, and explains the mechanism honestly: a single mistake weighs far more when the utterance is three words long, and there is no linguistic context to resolve ambiguity. What it does not do is divide.
Measured against each model's own headline figure on the random sample, every one of the six degrades by between 5.1 and 8.4 times on short dictation.
| Model | Real-world WER | Short dictation WER | Degradation |
|---|---|---|---|
| Wispr Canto | 3.4% | 21.4% | 6.3x |
| Google Gemini 3.1 Pro | 3.6% | 21.4% | 5.9x |
| OpenAI GPT Transcribe | 4.0% | 33.6% | 8.4x |
| AssemblyAI Universal-3.5 Pro | 4.2% | 21.4% | 5.1x |
| Google Gemini 3.5 Transcribe | 4.5% | 25.0% | 5.6x |
| Deepgram Nova-3 | 7.4% | 41.4% | 5.6x |
21.4% means roughly one word in five comes back wrong. Three different vendors deliver exactly that, so for short utterances the model is not the variable. Anything you dictate in bursts, a name into a field, a two word commit message, a quick reply, a slash command, sits in the flattest part of the curve, and switching providers will not move it. The one model that is clearly worse here is GPT Transcribe at 33.6%, which is a better-than-average model on the random sample and the second worst on short input.

The public benchmarks disagree with each other, and with both
Wispr also published results on three public English sets, which is where the launch narrative gets more interesting than the launch narrative claims.
| Benchmark | Canto | Gemini 3.1 Pro | Gemini 3.5 Transcribe | GPT Transcribe | AssemblyAI | Deepgram Nova-3 |
|---|---|---|---|---|---|---|
| FLEURS | 3.6% | 3.5% | 3.5% | 3.2% | 2.9% | 8.6% |
| LibriSpeech | 1.7% | 1.8% | 2.1% | 1.8% | 1.7% | 2.9% |
| Common Voice | 9.2% | 8.8% | 10.7% | 10.7% | 8.7% | 26.0% |
Track AssemblyAI Universal-3.5 Pro across all five measurements. It wins FLEURS outright, wins Common Voice outright and ties Canto for first on LibriSpeech. It then places fourth of six on real dictation and fifth of six on the challenge set. A buyer working from the public leaderboards would pick the model that finishes near the back on recorded speech, and a buyer working from Wispr's set would pick the one that ranks fifth of six on FLEURS. Both buyers would be reading real numbers.
There is a sharper version of the same problem inside our own archive. When we covered Gemini 3.5 Transcribe on 26 August, the FLEURS figures on record were 5.04% and 5.50% by way of Artificial Analysis. Wispr's chart puts the same model on the same named benchmark at 3.5%. Neither figure is wrong. Wispr states that it used the English subsets and evaluated without contextual prompting, and FLEURS is a multilingual set, so the two are almost certainly measuring different slices of it. That is our reading rather than a Wispr claim, and it is the point: the benchmark name is not the measurement, the configuration is. Our speech-to-text comparison and our GPT Transcribe tests carry the same caveat for the same reason.
What this evaluation does not prove
The 3.4% comes from a set nobody outside Wispr can run. It is 10 hours of that company's own product telemetry, and it is also, by construction, the distribution Canto was trained to fit. Wispr is explicit that it enforced speaker separation between the training and test sets to avoid overfitting on speaker characteristics. It does not claim domain separation, and there is no reason it would: the point of the model is to be good at exactly this. The competitors were not trained on it. A home-set advantage is not cheating, but a 0.2 percentage point margin measured on home turf is not a margin you should carry into your own decision.
The consent path is documented and worth reading before you enable anything. Wispr's data controls page states that if you allow data sharing, "your data (i.e. audio, transcript, edits)" may be used "to evaluate, train, or improve AI models, by Wispr," that enterprise accounts have sharing permanently disabled with no option to turn it on, and that the company does not sell data. Every sample in the 10 hour set came from a user who had that setting on.
The edits part is not incidental. Wispr describes using corrections as a training signal, identifying which edits are likely recognition errors using forced-alignment confidence and the shape of the edit, then grafting only that correction into the reference transcript used to score reinforcement learning rollouts. When you fix "cloud" to "Claude," that fix can become a label. When you rewrite the end of the sentence because you changed your mind, it should not, and the described pipeline tries to tell those apart.
Your custom dictionary can work against you
The most directly usable finding in the post is buried in the research section. At runtime, Canto is given specialized vocabulary from the user's dictionary, a standard technique in contextual ASR. Wispr measured what that costs. After reinforcement learning with GRPO, the model became more responsive to supplied context, including when the context was wrong: among the most phonetically similar terms, it adopted a false suggestion "nearly five times as often as the initial SFT model."
Their worked example is a good one. With "Barry" in your dictionary, "I want to make a berry salad" is at risk, and the risk grows in poor recording conditions where both are acoustically plausible. Wispr's fix during training was to feed phonetically similar distractors alongside the true term so the model has to select rather than copy, and it reports that this reduced false-context following while keeping correct-context adoption. The company is careful to say this does not yet define a final context policy.
For anyone using a dictation dictionary today, in any app, the operational lesson holds regardless of vendor: a large dictionary of names that sound like common words is not free. Prune entries that are near-homophones of words you actually say, and keep the list as short as the job allows.

What to do with this
Pick on the row that matches where you dictate, not the headline. If you work in a quiet room, the top four models sit inside 0.8 percentage points and you should be choosing on price, platform, latency and data policy instead, because the accuracy difference is not one you will perceive. If you dictate in noise, the gap is real and it is 3.64x wide, so this is the one case where switching models is the highest-leverage change available to you. If most of your input is one and two word bursts, stop shopping: the field ties at 21.4%, and the fix is workflow. Speak in full phrases rather than fragments, give the model a sentence of context to work with, and trim your custom dictionary of anything phonetically close to ordinary vocabulary.
Frequently asked questions
Is Canto available as an API?
No. The launch post describes Canto as the model powering real-time dictation in Wispr Flow and publishes no endpoint, no token pricing and no latency figure. Access is through the app's plans, which start at a free tier limited to 2,000 words a week on desktop.
Did Canto beat Gemini 3.1 Pro?
On Wispr's random sample of real dictations, yes, by 0.2 percentage points (3.4% against 3.6%). On the harder challenge set, no: Gemini 3.1 Pro leads at 8.2% against Canto's 9.1%, and it also wins the noisy subset at 6.1% against 7.6%. Wispr reports both losses in its own post and notes that Gemini 3.1 Pro is a much larger model not built for real-time use.
Why is the short dictation error rate so high for every model?
Two reasons, both given by Wispr. A single wrong word is a much larger share of a two word utterance than of a fifty word one, so WER rises mechanically. And a short utterance gives the model almost no surrounding language to disambiguate acoustically similar options. The result is 21.4% for the three best models, meaning roughly one word in five.
Can I reproduce these numbers?
Not the headline ones. The 10 hour real-world set and the 3 hour challenge set are private to Wispr, assembled from its own users' opted-in dictations. The three public benchmarks in the last table are reproducible in principle, though matching Wispr's exact conditions requires using the English subsets and running without contextual prompting, as its chart note specifies.
Does using Wispr Flow mean my dictations train the model?
Only if you turn the setting on. Wispr's data controls documentation says audio, transcripts and edits may be used to evaluate, train or improve its models when data sharing is enabled, managed under Settings, then Data and Privacy. Enterprise accounts have it permanently disabled.
Which model should I use for transcribing recorded audio rather than dictating?
These numbers do not answer that. Every set here is dictation, meaning short spontaneous speech captured live through laptop microphones, earbuds and headsets. Recorded interviews, podcasts and meetings are a different distribution with different requirements, especially diarization, which Wispr names as a target for its next model rather than a strength of this one.