Altworld released Hemmingway-1 on 20 September 2026 at 12:10 UTC: a 27B text model, Apache-2.0 weights, 262,144-token context, fine-tuned from Qwen3.8-27B. The model card leads with a first place. It reports 1026 on CommunicationBench against Fable 5.1 at 1024, and it says plainly that CommunicationBench is Altworld's own benchmark.

That disclosure is more honest than most launches manage. But the numbers behind it are published only inside chart images, not as text, so almost nobody has read them. We did. Across the five charts in the public repo, Hemmingway-1 tops two, places third on the one benchmark it does not own, and comes last of seven models on both story-writing columns of its own heatmap. The homepage asks "Why are you the best AI for creative writing?" The charts underneath it answer a narrower and more useful question.

What Altworld actually shipped

Hemmingway-1 is 27B parameters in bfloat16, 54.66 GB across twelve safetensors shards plus a 0.85 GB multi-token-prediction head. The config uses hybrid attention: linear-attention layers with a full-attention layer every fourth block, which is what makes a 262,144-token context tractable on a model this size. Weights are Apache-2.0, and so is the Qwen3.8-27B base they came from, so the licence chain is clean end to end. That is worth stating explicitly in a month when Qwen-Image-2.1 moved the other way and traded Apache 2.0 for a research licence.

Altworld is small. Its only prior public model is Astrea-R8-Chat-9B from July 2026, which sits at 10 likes. Hemmingway-1 passed 195 likes and 834 downloads in its first day, and the community shipped seventeen derivative repos in nineteen hours. Bartowski's GGUF set has been pulled 4,372 times, more than five times the original weights, which is the usual signal that the people trying a model are running it locally rather than reading the card.

There is also a product. Hemmingway.io runs the same model as a free web app with Mac, Windows and Android downloads, no iOS build, and paid tiers at $9, $29 and $79 a month, billed annually at $7.50, $24 and $65. So you can evaluate the model without touching the weights, which matters given what the weights weigh.

Two matte 3D bars engraved 834 and 4372 showing the community quant outpacing the official weights
First-day pulls: 834 for the official bf16 weights, 4,372 for the community GGUF build.

Five benchmarks, one of them independent

Four of the five charts carry an "INTERNAL" badge. Only EQ-Bench 4, the public emotional-intelligence leaderboard, does not, and Altworld says so in the fine print: "CommunicationBench, Human-Likeness and StoryBench are our own benchmarks. We built them, we ran them, and we are saying that up front." Here is every published score, read from the charts.

BenchmarkWhoseHemmingway-1RankLeaderGap
CommunicationBenchAltworld10261st of 9Fable 5.1, 1024+2
Human-LikenessAltworld10321st of 7Fable 5.1, 1006+26
EQ-Bench 4Public13303rd of 6Fable 5, 1341-11
StoryBenchAltworld11973rd-equal of 7Fable 5 Max, 1277-80
Message not memo (lower better)Altworld394th of 8GPT-6 Astra, 4+35

The methodology is defensible on its face. Eighty real requests, every answer matched head to head against another model's answer to the same request, shuffled so the judge could not see which was which, run in both orders so position could not sway the result, and judged by a model that was not among those being judged. Blind pairwise with order control is the right design. It is the scoreboard, not the protocol, that repays reading.

Two matte 3D bars of almost identical height engraved 1026 and 1024
CommunicationBench first and second place, two points apart on an eighty-prompt Elo.

The two-point win and the twenty-six-point win

CommunicationBench is the headline, and it is the weakest of the two internal wins. Hemmingway-1 scores 1026, Fable 5.1 scores 1024, Fable 5 scores 1009 and GLM-5.3 scores 1007. Four models inside nineteen points, with first and second separated by two, on an eighty-prompt Elo. That is a cluster, not a lead. The model card quotes the fifty-point gap over GPT-6 Astra at 976, which is real, and does not quote the two-point gap over the model actually in second.

Human-Likeness is the chart that earns the launch. Asked which of two texts a person wrote, judges took Hemmingway-1's at 1032 against Fable 5.1 at 1006, GLM-5.3 at 996, Fable 5 at 992, Kimi K3 at 985 and GPT-6 Astra at 964. Twenty-six points clear is a real separation, and the base model it was fine-tuned from sits at 952, so the training did something specific and measurable. If you are picking this model for one reason, that is the reason. It is not that it writes better. It is that it sounds less like a machine.

EQ-Bench 4, the one board Altworld does not own, puts it third at 1330 behind Fable 5 at 1341 and Kimi K3 at 1332, ahead of GPT-5.5 at 1316, Opus 4.7 at 1312 and Opus 4.8 at 1285. Eleven points off the top for a 27B model is a genuinely good result. Note what the chart does not contain, though: GPT-6 Astra, Grok 4.6 and Fable 5.1 all appear on the internal boards and none of them appear on this one. The independent comparison is against a different, older field than the internal comparison.

The story columns say something the homepage does not

Altworld tags the model "creative-writing" on Hugging Face and the site's own headline question is about being the best AI for creative writing. Its StoryBench chart puts Hemmingway-1 at 1197, level with Kimi K3, but 80 points behind Fable 5 Max at 1277 and 57 behind GLM-5.3 at 1254. Third-equal of seven.

The category heatmap is blunter. It scores human-likeness win share by request type, ties counting half, and Hemmingway-1's row reads: money and admin 81%, work 77%, hard asks 72%, explain and persuade 71%, personal 67%, public 55%, in voice 48%, story turns 27%, hostile storytelling 24%.

Those last two are the lowest numbers in their columns. Not lowest among the frontier models, lowest full stop: Fable 5.1 scores 81% on hostile storytelling and GLM-5.3 scores 76% on story turns, and even Gemini 3.8 Flash, which is last or near-last in every other column on the chart, beats Hemmingway-1 on both at 29% and 33%. Altworld says as much in prose, that it loses on hostile storytelling and long story turns and that the story models are better at those. The prose is accurate. The tag and the marketing question are the things that do not match the chart.

The other number worth your attention is "in voice" at 48%. That is the category where you hand it a rough note and ask it to sound like you, which the site calls its favourite job. Judges preferred its text less than half the time, and it places fourth there behind Kimi K3 at 61%, GLM-5.3 at 58% and GPT-6 Astra at 56%. The model that wins "sounds like a person" does not win "sounds like this particular person," and for anyone with an established voice to protect, those are different products.

Three matte 3D bars in true proportion engraved 81, 27 and 24
Win share by category: 81% on money and admin, 27% on story turns, 24% on hostile storytelling.

It wraps the message more often than the model it came from

The fifth chart is titled "The Message, Not a Memo" and measures how often a model buries the actual text in commentary, options and notes you have to read past. Lower is better. The section above it in the model card says you get the message, not a memo, and names Fable 5, GLM-5.3 and Kimi K3 as burying it in more than nine replies out of ten, which the chart supports at 93%, 93% and 92%.

The chart also shows GPT-6 Astra at 4%, Qwen3.8-27B base at 19% and Grok 4.6 at 22%, all of them cleaner than Hemmingway-1 at 39%. Three models hand over the bare text more reliably than it does, and one of them is the model it was fine-tuned from. Whatever the fine-tune bought in human-likeness, it cost twenty points on not padding the answer. None of those three are named in the paragraph above the chart. This is the same pattern we found when Wispr Canto published its error rates only as chart images: the prose selects, the chart discloses, and the chart is the thing worth opening.

Three matte 3D bars in true proportion engraved 4, 19 and 39 where taller is worse
How often the answer is buried in commentary: GPT-6 Astra 4%, the Qwen base 19%, Hemmingway-1 39%.

What it costs to run

Full bf16 weights are 54.66 GB, which means two 40 GB accelerators or one 80 GB card before you account for the KV cache at a 262,144-token context. Community quantisation landed the same day and is what most people will actually use.

BuildSizeFits
bf16 safetensors (official)54.66 GB80 GB card, or 2x 40 GB
Q8_0 GGUF29.12 GB32 GB card, 48 GB Mac
Q6_K GGUF23.86 GB24 GB card, tight
Q4_K_M GGUF17.44 GB24 GB card, comfortable
IQ4_XS GGUF15.48 GB16 GB card, tight
Q3_K_M GGUF13.40 GB16 GB card
IQ2_XXS GGUF8.88 GB12 GB card

The served path is one line, vllm serve Altworld/Hemmingway-1 --max-model-len 262144, documented against vLLM, and there is a standard Transformers snippet for a single-shot call. Apple silicon users have 4-bit and 8-bit MLX conversions from day one. Remember that a 262,144-token context is a ceiling, not a free allowance: the KV cache at full length will cost you more memory than the quantised weights do.

Who this is actually for

Take it if you write a high volume of short, real correspondence and you want it to sound like you wrote it: client emails, follow-ups, the awkward note, release notes, replies you have been avoiding. Money and admin at 81%, work at 77% and hard asks at 72% are strong, consistent numbers, the 26-point human-likeness margin is the one result on the page with clear daylight in it, and Apache-2.0 means you can run it commercially without asking anyone.

Skip it if you are writing fiction, scripts, long narrative or anything with a sustained character voice. Its maker's own charts put it third-equal on StoryBench and last of seven on both story columns, and the model card agrees. Skip it too if your test is "write this in my voice," where 48% means a coin flip. And do not choose it on the CommunicationBench first place alone, because two points on an eighty-prompt internal Elo is not a difference you will feel. Pull the Q4_K_M build, run twenty of your own real messages through it beside whatever you use now, and judge the human-likeness claim yourself. It is the claim the charts actually support.

Frequently asked questions

Is Hemmingway-1 really open source?

The weights are Apache-2.0 and usable commercially, and the base model Qwen3.8-27B carries the same licence, so nothing upstream restricts you. The training data and the benchmark harnesses are not published, so it is open weights rather than fully open source. The evaluation code for CommunicationBench, Human-Likeness and StoryBench is not available, which means those three results cannot be independently reproduced.

Which of its benchmarks can I trust?

EQ-Bench 4 is the only one Altworld does not own, and it places the model third of the six charted at 1330. The other four are internal. That does not make them wrong, and the blind, both-orders, separate-judge protocol is a reasonable design, but they are unreproducible and the model's maker chose the field, the prompts and the judge.

Can it write stories?

It can, but it is not what you should pick for the job. On Altworld's own StoryBench it scores 1197 against Fable 5 Max at 1277, and on its own category heatmap it scores 24% on hostile storytelling and 27% on long story turns, the lowest figures in both columns. The model card states outright that story models are better at those.

What hardware do I need?

Full precision is 54.66 GB, so an 80 GB accelerator or two 40 GB cards. The Q4_K_M GGUF at 17.44 GB runs comfortably on a 24 GB consumer card, IQ4_XS at 15.48 GB fits 16 GB, and IQ2_XXS at 8.88 GB will load on 12 GB if you accept the quality loss. Budget extra for the KV cache if you use the long context.

Do I have to download the weights to try it?

No. Hemmingway.io runs the same model as a free web app with a usage-limited tier, plus Mac, Windows and Android clients and an API. There is no iOS app. Paid plans are $9, $29 and $79 a month, or $7.50, $24 and $65 billed annually.

How does it compare to its base model?

Against Qwen3.8-27B it gains 72 points on CommunicationBench (1026 against 954), 80 on Human-Likeness (1032 against 952) and 504 on StoryBench (1197 against 693). It also regresses on one measure: it wraps its answer in commentary 39% of the time against the base model's 19%.